Introduction to Psychology · Memory and Cognition

Intelligence Theories and Intelligence Testing

9 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 7 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Study tools

In 30 seconds

is the general capacity to learn, reason, and adapt, though theorists disagree about whether it is one thing (Spearman's g), a pair of broad abilities (), a triad (Sternberg), or many independent strengths (Gardner). Modern IQ tests — descendants of 's scale, now the , , and — compare a person's score to age-based . Good tests are standardized, reliable, and valid, but every test faces questions of and bias, and heritability estimates describe group variation, not any individual's fixed potential.

Why this matters

Intelligence tests are used in educational placement, learning-disability assessment, and some clinical and vocational evaluations — but always as one piece of evidence, never as a label or ceiling. Responsible practice pairs test scores with interviews, classroom observations, and adaptive-behavior measures, and interprets them in light of language background, culture, and opportunity. The concept of matters directly here: a score depressed by an unfamiliar language or testing format can misrepresent ability. Ethical use therefore emphasizes cultural fairness, avoids deterministic labeling of individuals, and treats any single IQ number as a snapshot with measurement error, not a verdict on a person's worth or future. This is educational information about testing practice, not an assessment or diagnosis of any individual.

The college version

1. Theories of Intelligence

Intelligence is commonly defined as the ability to learn from experience, solve problems, and adapt to the environment. Theories differ on its structure:

  • General intelligence (Spearman's g): Charles Spearman found that scores on different mental tasks correlate positively and proposed a single underlying factor, g, that influences performance across all cognitive tasks.
  • Fluid vs. crystallized intelligence (Cattell and Horn): Fluid intelligence is the capacity to reason quickly and solve novel problems independent of prior knowledge (peaks in early adulthood); crystallized intelligence is accumulated knowledge, vocabulary, and skills (grows across the lifespan).
  • theory: Intelligence has three aspects — analytical (problem solving and academic skills), creative (generating novel ideas), and practical ("street smarts," adapting to everyday contexts).
  • : Intelligence is not one ability but at least eight relatively independent intelligences (e.g., linguistic, logical-mathematical, spatial, musical, bodily-kinesthetic, interpersonal, intrapersonal, naturalist). The theory is popular in education but criticized for weak empirical support and unclear boundaries with talent or personality.

2. History and Tools of Intelligence Testing

  • Alfred Binet (early 1900s) created the first practical intelligence test to identify schoolchildren needing extra help, measuring mental age. He warned against viewing the score as a fixed, inborn quantity.
  • Stanford-Binet: Lewis Terman's American adaptation introduced the intelligence quotient (IQ), originally the ratio of mental age to chronological age times 100.
  • Wechsler scales: The WAIS (Wechsler Adult Intelligence Scale) and WISC (Wechsler Intelligence Scale for Children) are the most widely used modern tests. They yield an overall IQ plus index scores (e.g., verbal comprehension, working memory, perceptual reasoning, processing speed).
  • IQ: In modern tests, IQ is a standardized score with a mean of 100 and a standard deviation of about 15, so roughly two-thirds of people score between 85 and 115.

3. Test Quality and Fairness

A good test must be standardized — administered and scored the same way for everyone, with norms (scores of a large, representative comparison group) that let an individual's score be interpreted relative to peers. Quality is judged by (consistency of scores across time and forms) and validity (whether the test measures what it claims), including predictive validity (how well it forecasts relevant outcomes such as grades or job performance). Cultural fairness and test bias are persistent concerns: if items or norms favor one cultural or socioeconomic group, the test may measure opportunity and familiarity as much as ability. A test can be statistically reliable yet still produce group differences driven by bias in item content, norming samples, or testing conditions — a reminder that "valid" is always relative to a purpose and population.

How it works

  1. Test authors define the construct (intelligence) and write many candidate items.
  2. The test is standardized on a large, representative sample to build norms.
  3. An individual's raw score is converted to an IQ relative to same-age peers.
  4. Psychometricians evaluate reliability and validity, including predictive validity.
  5. Users interpret scores cautiously, attending to cultural fairness and possible test bias.
  6. Conclusions about nature/nurture respect the limits of heritability and avoid treating any score as fixed potential.

Common confusions

Do not confuseWithDifference
ReliabilityValidityReliability = consistency; validity = measuring the right thing.
Fluid intelligenceCrystallized intelligenceFluid = novel reasoning; crystallized = accumulated knowledge.
Spearman's gGardner's multiple intelligencesOne general factor vs. several independent abilities.
StandardizationNormsStandardization = the uniform procedure; norms = the comparison-group scores.
HeritabilityInnate/fixed for an individualHeritability is a population-level variance ratio, not individual destiny.
Test biasMean group differenceBias is systematic measurement error; a difference may reflect real or environmental factors.
IQAchievementIQ aims at general capacity; achievement measures what has been learned.

Memory aids

Remember "Gardner Grows Many Skills; Spearman Sees One g" for the theories, and "S-R-V" for test quality — Standardized and Normed, then judged by Reliability and Validity. For the WAIS/WISC pair, recall "A for Adult, C for Child."

Quick review

Topic Recap

Intelligence is theorized as a single general factor (g), as fluid vs. crystallized abilities, as Sternberg's analytical/creative/practical triad, or as Gardner's multiple intelligences. Modern tests (Stanford-Binet, WAIS, WISC) standardize scores into an IQ with mean 100 and SD 15, and their worth depends on standardization, norms, reliability, and validity — including predictive validity. Cultural fairness and test bias are unresolved concerns, and the nature/nurture question must respect the limits of heritability.

Knowledge Check

  1. Which theory proposes a single general factor, g, underlying all cognitive abilities?
  2. A test gives the same score when taken twice. Which quality does this reflect?
  3. The ability to solve novel problems independent of prior knowledge is called what?
  4. Which test is used to measure intelligence in adults?
  5. Why must heritability estimates be interpreted with caution?

Answers and Rationales

  1. Spearman's general intelligence — factor analysis showed positive correlations across tasks, suggesting one underlying factor.
  2. Reliability — consistency across administrations is the definition of reliability.
  3. Fluid intelligence — novel, knowledge-independent reasoning; crystallized intelligence is accumulated knowledge.
  4. WAIS (Wechsler Adult Intelligence Scale) — the WISC is the children's version.
  5. Because heritability is a population statistic about variance, not a statement about any individual's fixed potential — it also changes with the environment and cannot explain between-group differences.
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Think of intelligence as a car's overall performance. Spearman's g is the claim that there is one "engine quality" powering everything the car does. The fluid vs. crystallized view says there are really two systems: fluid intelligence is the engine's raw power for novel problems (fast, peaks early), while crystallized intelligence is the accumulated fuel of knowledge and skills (grows with experience). Gardner says a car isn't one thing at all — it's the engine and the brakes and the sound system, each a separate talent. Sternberg says what matters is how well the car handles three real jobs: analyzing, creating, and getting around in the real world. An IQ test is a mechanic's inspection: it scores your car against the average car of the same age, giving a number with a mean of 100.

Where it stops being exact: a real car's performance is measurable with a stopwatch; intelligence is inferred from test scores, which sample only some abilities. The car analogy also hides the fact that intelligence tests are influenced by culture, schooling, and opportunity — and that a single number cannot capture a person's full potential.

Simple Example

Two students both score 110 on an IQ test. One is a 10-year-old and one is an adult — the 10-year-old's raw score was compared to other 10-year-olds, not to the adult. This is what standardization against age-based norms means: the IQ number is always relative to a comparison group, not an absolute quantity.

Worked example

  1. Spearman's factor analysis (1904): Correlations among school subjects and mental tests led Spearman to posit a single g factor. Modern hierarchical models retain a general factor above more specific abilities — evidence that performance across tasks is correlated, though the interpretation of g remains debated.
  2. The Flynn effect: Average IQ scores rose substantially across the 20th century, showing that IQ is responsive to environment (nutrition, schooling, complexity of life) and undermining any claim that IQ is fixed. This is a strong argument for environmental influence.
  3. Nature/nurture and heritability: Twin and adoption studies estimate heritability — the proportion of variance in a trait within a particular population attributable to genetic differences. Estimates for IQ in studied populations are substantial, but the limits of heritability are critical: heritability is a population statistic, not a statement about any one person; it says nothing about a specific individual's potential; it can change when environments change; and high heritability within a group does not explain average differences between groups.
  4. Methodological limits: Much evidence is correlational. A correlation between IQ and school grades (predictive validity) does not show IQ causes achievement — other factors (motivation, teaching quality, opportunity) may drive both. Cross-group comparisons are especially fraught: differences can reflect test bias, unequal opportunity, stereotype threat, and norming problems rather than ability. Cultural fairness is therefore an empirical question answered case by case, not an assumption. Gardner's and Sternberg's theories, while influential, have less standardized measurement support than g-based models.

Key takeaways

  • High yield: Spearman's g = one general factor; fluid vs. crystallized = two broad abilities (novel reasoning vs. accumulated knowledge).
  • High yield: Modern IQ has a mean of 100 and SD of ~15; it is relative to age-based norms, not absolute.
  • High yield: Reliability = consistency; validity = accuracy/measuring the right thing; predictive validity = forecasting outcomes.
  • High yield: Standardization + norms are what make an IQ score interpretable.
  • Heritability describes variance within a population, never an individual's fixed potential; the Flynn effect shows IQ responds to environment.
  • Gardner's theory is popular but empirically weak; Sternberg adds creative and practical intelligences.
  • Correlations between IQ and outcomes do not prove causation.
  • Test bias and cultural fairness must be evaluated empirically for each test and population.

Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsYou’ll learn to · Key vocabulary · Related

You’ll learn to

  • Define intelligence and compare Spearman's g with the fluid vs. crystallized distinction, Sternberg's triarchic theory, and Gardner's multiple intelligences.
  • Trace the history of testing from Binet to the Stanford-Binet, WAIS, and WISC, and explain IQ.
  • Explain how tests are developed through standardization and norms, and evaluate them using reliability, validity, and predictive validity.
  • Discuss cultural fairness, test bias, and the nature/nurture debate, including the limits of heritability.

Key vocabulary

Intelligence
The ability to learn, reason, and adapt.
General intelligence (g)
Spearman's single underlying cognitive factor.
Fluid vs. crystallized
Novel-reasoning ability vs. accumulated knowledge.
Sternberg's triarchic
Analytical, creative, and practical intelligences.
Gardner's multiple intelligences
Several relatively independent intelligences.
Binet
Creator of the first practical intelligence test.
Stanford-Binet
Terman's adaptation that introduced IQ.
WAIS
Wechsler Adult Intelligence Scale.
WISC
Wechsler Intelligence Scale for Children.
IQ
Standardized score with mean 100, SD 15.
Standardization
Uniform administration and scoring.
Norms
Scores of a representative comparison group.
Reliability
Consistency of measurement.
Validity
Whether a test measures what it claims.
Predictive validity
How well a test forecasts outcomes.
Cultural fairness
Whether a test is fair across cultural groups.
Test bias
Systematic error favoring one group.
Nature/nurture
Genes vs. environment in shaping a trait.
Heritability limits
Heritability is a population, not individual, statistic.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.