Introduction to Psychology · Learning

Operant Conditioning

6 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 7 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Study tools

In 30 seconds

is learning in which the consequences of a behavior change how likely that behavior is to recur. 's states that behaviors followed by satisfying outcomes become more likely, while those followed by unpleasant outcomes become less likely. B. F. systematized these ideas, showing that strengthens behavior and weakens it. Complex behavior is built through and maintained by schedules of reinforcement.

Why this matters

Operant principles are widely used in education and health behavior. Token economies — in which students or patients earn secondary reinforcers (points, tokens) exchangeable for privileges — apply to encourage target behaviors. Behavior-change programs for habits like taking medication or exercising often rely on shaping (starting small) and intermittent reinforcement. also explains avoidance patterns, such as procrastination, which is reinforced because it temporarily removes the anxiety of a task. This is educational content about learning principles, not clinical treatment advice.

The college version

1. Thorndike and the Law of Effect

Edward Thorndike placed cats in "puzzle boxes" and observed them learn to escape by trial and error. He proposed the law of effect: responses that produce a satisfying effect become more likely to recur, and responses that produce a discomforting effect become less likely. This principle is the foundation of operant conditioning.

2. Reinforcement and Punishment

Reinforcement is any consequence that increases a behavior. Positive reinforcement adds a desirable stimulus (a treat, praise, money). Negative reinforcement removes an unpleasant stimulus (turning off a nagging alarm by getting up). Both increase behavior; "negative" means something is taken away, not that it is bad. Punishment is any consequence that decreases behavior. adds something unpleasant (a scolding, a fine). removes something desirable (losing screen time). A satisfies a biological need (food, water); a secondary (conditioned) reinforcer gains value through association (money, grades).

3. Shaping and Schedules of Reinforcement

Shaping reinforces successive approximations — small steps toward the target behavior. Extinction occurs when a previously reinforced behavior stops producing the reinforcer, so it declines. Continuous reinforcement reinforces every correct response: fast learning, rapid extinction. Partial (intermittent) reinforcement reinforces only some responses, producing slower extinction. The four partial schedules are fixed-ratio (after a set number of responses), variable-ratio (after an unpredictable number), fixed-interval (first response after a set time), and variable-interval (first response after an unpredictable time).

How it works

  1. A behavior occurs, spontaneously or prompted.
  2. It is followed by a consequence — a reinforcer or a punisher.
  3. If reinforced, the behavior becomes more frequent; if punished, less frequent.
  4. Complex behaviors are built by reinforcing successive approximations (shaping).
  5. The behavior is maintained by a schedule — continuous or partial.
  6. If reinforcement stops entirely, the behavior undergoes extinction.

Common confusions

Do not confuseWithDifference
Negative reinforcementPunishmentNegative reinforcement increases behavior by removing; punishment decreases behavior
Positive punishmentNegative reinforcementPositive punishment adds unpleasant to decrease; negative reinforcement removes unpleasant to increase
Operant conditioningClassical conditioningOperant shapes voluntary behavior via consequences; classical pairs stimuli for involuntary responses
Continuous reinforcementPartial reinforcementContinuous reinforces every response (fast extinction); partial reinforces some (slow extinction)
Fixed-intervalFixed-ratioFixed-interval depends on time; fixed-ratio on number of responses

Memory aids

Remember the four consequence types with "ARP-N": Add good (positive reinforcement), Remove bad (negative reinforcement), Add bad (positive punishment), Remove good (negative punishment) — Note the two reinforcements increase behavior and the two punishments decrease it. For schedules, think "R = responses, I = time": ratio counts responses, interval counts time; fixed is predictable, variable is unpredictable.

Quick review

Topic Recap

Operant conditioning explains how consequences shape voluntary behavior. Thorndike's law of effect, systematized by Skinner, shows that reinforcement increases behavior and punishment decreases it. Positive and negative refer to whether a stimulus is added or removed, not whether the outcome is pleasant. Primary reinforcers satisfy biological needs; secondary reinforcers are learned. Shaping builds complex behavior through successive approximations, and schedules of reinforcement determine how quickly behavior is learned and how resistant it is to extinction.

Knowledge Check

  1. The law of effect was proposed by ______.
  2. Giving a child a sticker for completing homework is ______ reinforcement.
  3. A rat learns to press a lever to stop an ongoing shock. This is ______ reinforcement.
  4. Which schedule produces the highest, steadiest response rate: fixed-interval or variable-ratio?
  5. True or False: Punishment is the most reliable way to teach a new desired behavior.

Answers and Rationales

  1. Thorndike. His law of effect stated that satisfying consequences stamp behaviors in and unpleasant consequences stamp them out.
  2. Positive. A desirable stimulus (the sticker) is added to increase homework completion.
  3. Negative. The behavior (pressing the lever) removes an unpleasant stimulus (the shock), increasing lever pressing.
  4. Variable-ratio. Reinforcement arrives after an unpredictable number of responses, so responding is fast, steady, and resists extinction.
  5. False. Punishment suppresses behavior but does not teach what to do instead; reinforcement is more effective for building new behaviors.
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Imagine teaching a dog to sit. You do not explain the concept — you wait until its bottom nears the floor, say "sit," and give a treat. The treat makes the dog more likely to sit next time. The dog is not reasoning; it simply does more of what pays off and less of what does not.

That is operant conditioning: behavior is controlled by its consequences. Reinforcement (something good follows) increases a behavior; punishment (something bad follows) decreases it. This differs from classical conditioning (Topic 15), where an automatic response is attached to a new trigger. Here the learner acts voluntarily and learns from results.

Where it stops being exact: the "treat" story suggests consequences always work instantly and mechanically. In reality, timing matters, the learner's biology and history matter, and reinforcement works even when occasional. Operant conditioning also under-explains insight, reasoning, and learning without reward (Topic 17).

Simple Example

A student answers a question correctly and the teacher praises her. She raises her hand more often afterward. The praise (a reinforcer) followed the behavior (answering), making it more frequent.

Worked example

  1. Skinner extended Thorndike's law with controlled experiments using the "operant chamber" (Skinner box), where a rat or pigeon could press a lever or peck a key for food, allowing precise measurement.
  2. Schedules produce characteristic patterns: variable-ratio schedules generate very high, steady response rates (the logic behind slot machines), while fixed-interval schedules produce a pause after reinforcement then accelerating responding.
  3. Correlation vs. causation: studies linking heavy social-media use to reward-seeking are usually correlational and cannot prove a schedule caused the behavior, because third variables (personality, environment) may be involved.
  4. Methodological limits: much operant work used nonhuman animals under controlled conditions; extrapolating to humans requires caution, since human behavior is also shaped by language, thought, and social norms (and instinctive drift).
  5. Limitations of the theory: operant conditioning under-explains learning without reinforcement, insight, and observational learning (Topic 17); modern accounts treat it as one mechanism among several.

Key takeaways

  • High yield: Reinforcement increases behavior; punishment decreases it — focus on the effect, not the intent.
  • High yield: "Positive" means adding a stimulus; "negative" means removing one.
  • High yield: Negative reinforcement strengthens behavior; it is not punishment.
  • Variable-ratio schedules produce the highest, steadiest response rates.
  • High yield: Partial reinforcement produces slower extinction than continuous reinforcement.
  • High yield: Operant conditioning explains voluntary behavior; classical conditioning explains involuntary responses.

Keep learning

Ready to build on this? Continue to the next lesson.

Study toolsYou’ll learn to · Key vocabulary

You’ll learn to

  • Define operant conditioning and explain the law of effect.
  • Distinguish positive and negative reinforcement from positive and negative punishment.
  • Describe shaping and the four schedules of reinforcement (fixed-ratio, variable-ratio, fixed-interval, variable-interval).
  • Identify primary versus secondary reinforcers and the limits of operant explanations.

Key vocabulary

Operant conditioning
Learning controlled by consequences
Thorndike
Psychologist who studied trial-and-error learning in cats
Law of effect
Satisfying outcomes repeat behavior; unpleasant ones do not
Skinner
Psychologist who systematized operant conditioning
Reinforcement
Any consequence that increases a behavior
Positive reinforcement
Adding a desirable stimulus to increase behavior
Negative reinforcement
Removing an unpleasant stimulus to increase behavior
Primary reinforcer
Reinforcer satisfying a biological need
Secondary reinforcer
Reinforcer gaining value through association
Punishment
Any consequence that decreases a behavior
Positive punishment
Adding something unpleasant to decrease behavior
Negative punishment
Removing something desirable to decrease behavior
Shaping
Reinforcing successive approximations toward a target
Successive approximations
Small steps toward the final behavior
Extinction
Decline when reinforcement stops
Continuous reinforcement
Reinforcing every correct response
Partial reinforcement
Reinforcing only some responses
Fixed-ratio / variable-ratio / fixed-interval / variable-interval
Schedules by responses or time, fixed or variable

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.