Introduction to Psychology · Learning
Operant Conditioning
On this page 7 sections
In 30 seconds
Operant conditioning Learning controlled by consequences Full entry → is learning in which the consequences of a behavior change how likely that behavior is to recur. Thorndike Psychologist who studied trial-and-error learning in cats Full entry →'s Law of effect Satisfying outcomes repeat behavior; unpleasant ones do not Full entry → states that behaviors followed by satisfying outcomes become more likely, while those followed by unpleasant outcomes become less likely. B. F. Skinner Psychologist who systematized operant conditioning Full entry → systematized these ideas, showing that Reinforcement Any consequence that increases a behavior Full entry → strengthens behavior and Punishment Any consequence that decreases a behavior Full entry → weakens it. Complex behavior is built through Shaping Reinforcing successive approximations toward a target Full entry → and maintained by schedules of reinforcement.
Why this matters
Operant principles are widely used in education and health behavior. Token economies — in which students or patients earn secondary reinforcers (points, tokens) exchangeable for privileges — apply Positive reinforcement Adding a desirable stimulus to increase behavior Full entry → to encourage target behaviors. Behavior-change programs for habits like taking medication or exercising often rely on shaping (starting small) and intermittent reinforcement. Negative reinforcement Removing an unpleasant stimulus to increase behavior Full entry → also explains avoidance patterns, such as procrastination, which is reinforced because it temporarily removes the anxiety of a task. This is educational content about learning principles, not clinical treatment advice.
The college version
1. Thorndike and the Law of Effect
Edward Thorndike placed cats in "puzzle boxes" and observed them learn to escape by trial and error. He proposed the law of effect: responses that produce a satisfying effect become more likely to recur, and responses that produce a discomforting effect become less likely. This principle is the foundation of operant conditioning.
2. Reinforcement and Punishment
Reinforcement is any consequence that increases a behavior. Positive reinforcement adds a desirable stimulus (a treat, praise, money). Negative reinforcement removes an unpleasant stimulus (turning off a nagging alarm by getting up). Both increase behavior; "negative" means something is taken away, not that it is bad. Punishment is any consequence that decreases behavior. Positive punishment Adding something unpleasant to decrease behavior Full entry → adds something unpleasant (a scolding, a fine). Negative punishment Removing something desirable to decrease behavior Full entry → removes something desirable (losing screen time). A Primary reinforcer Reinforcer satisfying a biological need Full entry → satisfies a biological need (food, water); a secondary (conditioned) reinforcer gains value through association (money, grades).
3. Shaping and Schedules of Reinforcement
Shaping reinforces successive approximations — small steps toward the target behavior. Extinction occurs when a previously reinforced behavior stops producing the reinforcer, so it declines. Continuous reinforcement reinforces every correct response: fast learning, rapid extinction. Partial (intermittent) reinforcement reinforces only some responses, producing slower extinction. The four partial schedules are fixed-ratio (after a set number of responses), variable-ratio (after an unpredictable number), fixed-interval (first response after a set time), and variable-interval (first response after an unpredictable time).
How it works
- A behavior occurs, spontaneously or prompted.
- It is followed by a consequence — a reinforcer or a punisher.
- If reinforced, the behavior becomes more frequent; if punished, less frequent.
- Complex behaviors are built by reinforcing successive approximations (shaping).
- The behavior is maintained by a schedule — continuous or partial.
- If reinforcement stops entirely, the behavior undergoes extinction.
Common confusions
| Do not confuse | With | Difference |
|---|---|---|
| Negative reinforcement | Punishment | Negative reinforcement increases behavior by removing; punishment decreases behavior |
| Positive punishment | Negative reinforcement | Positive punishment adds unpleasant to decrease; negative reinforcement removes unpleasant to increase |
| Operant conditioning | Classical conditioning | Operant shapes voluntary behavior via consequences; classical pairs stimuli for involuntary responses |
| Continuous reinforcement | Partial reinforcement | Continuous reinforces every response (fast extinction); partial reinforces some (slow extinction) |
| Fixed-interval | Fixed-ratio | Fixed-interval depends on time; fixed-ratio on number of responses |
Memory aids
Remember the four consequence types with "ARP-N": Add good (positive reinforcement), Remove bad (negative reinforcement), Add bad (positive punishment), Remove good (negative punishment) — Note the two reinforcements increase behavior and the two punishments decrease it. For schedules, think "R = responses, I = time": ratio counts responses, interval counts time; fixed is predictable, variable is unpredictable.
Quick review
Topic Recap
Operant conditioning explains how consequences shape voluntary behavior. Thorndike's law of effect, systematized by Skinner, shows that reinforcement increases behavior and punishment decreases it. Positive and negative refer to whether a stimulus is added or removed, not whether the outcome is pleasant. Primary reinforcers satisfy biological needs; secondary reinforcers are learned. Shaping builds complex behavior through successive approximations, and schedules of reinforcement determine how quickly behavior is learned and how resistant it is to extinction.
Knowledge Check
- The law of effect was proposed by ______.
- Giving a child a sticker for completing homework is ______ reinforcement.
- A rat learns to press a lever to stop an ongoing shock. This is ______ reinforcement.
- Which schedule produces the highest, steadiest response rate: fixed-interval or variable-ratio?
- True or False: Punishment is the most reliable way to teach a new desired behavior.
Answers and Rationales
- Thorndike. His law of effect stated that satisfying consequences stamp behaviors in and unpleasant consequences stamp them out.
- Positive. A desirable stimulus (the sticker) is added to increase homework completion.
- Negative. The behavior (pressing the lever) removes an unpleasant stimulus (the shock), increasing lever pressing.
- Variable-ratio. Reinforcement arrives after an unpredictable number of responses, so responding is fast, steady, and resists extinction.
- False. Punishment suppresses behavior but does not teach what to do instead; reinforcement is more effective for building new behaviors.

Eli explains
The same idea, in plain words
Explain it like I’m 10
Imagine teaching a dog to sit. You do not explain the concept — you wait until its bottom nears the floor, say "sit," and give a treat. The treat makes the dog more likely to sit next time. The dog is not reasoning; it simply does more of what pays off and less of what does not.
That is operant conditioning: behavior is controlled by its consequences. Reinforcement (something good follows) increases a behavior; punishment (something bad follows) decreases it. This differs from classical conditioning (Topic 15), where an automatic response is attached to a new trigger. Here the learner acts voluntarily and learns from results.
Where it stops being exact: the "treat" story suggests consequences always work instantly and mechanically. In reality, timing matters, the learner's biology and history matter, and reinforcement works even when occasional. Operant conditioning also under-explains insight, reasoning, and learning without reward (Topic 17).
Simple Example
A student answers a question correctly and the teacher praises her. She raises her hand more often afterward. The praise (a reinforcer) followed the behavior (answering), making it more frequent.
Worked example
- Skinner extended Thorndike's law with controlled experiments using the "operant chamber" (Skinner box), where a rat or pigeon could press a lever or peck a key for food, allowing precise measurement.
- Schedules produce characteristic patterns: variable-ratio schedules generate very high, steady response rates (the logic behind slot machines), while fixed-interval schedules produce a pause after reinforcement then accelerating responding.
- Correlation vs. causation: studies linking heavy social-media use to reward-seeking are usually correlational and cannot prove a schedule caused the behavior, because third variables (personality, environment) may be involved.
- Methodological limits: much operant work used nonhuman animals under controlled conditions; extrapolating to humans requires caution, since human behavior is also shaped by language, thought, and social norms (and instinctive drift).
- Limitations of the theory: operant conditioning under-explains learning without reinforcement, insight, and observational learning (Topic 17); modern accounts treat it as one mechanism among several.
Key takeaways
- High yield: Reinforcement increases behavior; punishment decreases it — focus on the effect, not the intent.
- High yield: "Positive" means adding a stimulus; "negative" means removing one.
- High yield: Negative reinforcement strengthens behavior; it is not punishment.
- Variable-ratio schedules produce the highest, steadiest response rates.
- High yield: Partial reinforcement produces slower extinction than continuous reinforcement.
- High yield: Operant conditioning explains voluntary behavior; classical conditioning explains involuntary responses.
Study toolsYou’ll learn to · Key vocabulary
You’ll learn to
- Define operant conditioning and explain the law of effect.
- Distinguish positive and negative reinforcement from positive and negative punishment.
- Describe shaping and the four schedules of reinforcement (fixed-ratio, variable-ratio, fixed-interval, variable-interval).
- Identify primary versus secondary reinforcers and the limits of operant explanations.
Key vocabulary
- Operant conditioning
- Learning controlled by consequences
- Thorndike
- Psychologist who studied trial-and-error learning in cats
- Law of effect
- Satisfying outcomes repeat behavior; unpleasant ones do not
- Skinner
- Psychologist who systematized operant conditioning
- Reinforcement
- Any consequence that increases a behavior
- Positive reinforcement
- Adding a desirable stimulus to increase behavior
- Negative reinforcement
- Removing an unpleasant stimulus to increase behavior
- Primary reinforcer
- Reinforcer satisfying a biological need
- Secondary reinforcer
- Reinforcer gaining value through association
- Punishment
- Any consequence that decreases a behavior
- Positive punishment
- Adding something unpleasant to decrease behavior
- Negative punishment
- Removing something desirable to decrease behavior
- Shaping
- Reinforcing successive approximations toward a target
- Successive approximations
- Small steps toward the final behavior
- Extinction
- Decline when reinforcement stops
- Continuous reinforcement
- Reinforcing every correct response
- Partial reinforcement
- Reinforcing only some responses
- Fixed-ratio / variable-ratio / fixed-interval / variable-interval
- Schedules by responses or time, fixed or variable
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.
