Population Health for Nurses · Evidence-Based Decision-Making
Evaluating the Quality of the Evidence
On this page 9 sections
In 30 seconds
Finding evidence (Topic 2) is half the job; the other half is deciding how much to trust it. Critical appraisal Systematically judging whether a study is valid and applicable Full entry → is the disciplined habit of asking whether a study was conducted well enough to believe, and whether its findings apply to the community you serve. A published article is a claim, not a proof: quality depends on design, conduct, and reporting, not on the journal's reputation. Appraisal asks two questions, in order: Is the finding valid? (Internal validity Whether the study's result is true for the people studied Full entry →) and Will it work here? (External validity Whether the result transfers to other populations and settings Full entry →). Both must pass before an intervention is worth adopting.
Why this matters
- Weak evidence causes real harm. An unproven program wastes public money, delays effective help, and erodes community trust.
- Bias Systematic error that distorts results (selection, information) Full entry → is invisible in the summary. A press release or an abstract can look convincing while hiding fatal flaws (no comparison group, high dropout, cherry-picked outcomes).
- Population health decisions compound. A program adopted at county scale affects thousands of people — a small error in judgment becomes a large error in reach.
- Professional credibility and exams. Nurses who can say why a study is trustworthy are taken seriously in policy settings; evidence hierarchies, validity, bias, and Confounding A third variable tangled with exposure and outcome Full entry → are staple test items.
The college version
Core Concepts
Study design and the evidence hierarchy
For questions about whether an intervention works, study designs form a rough hierarchy from strongest to weakest:
- Systematic reviews and meta-analyses — transparent syntheses of multiple studies.
- Randomized controlled trials (RCTs) — participants randomly assigned to intervention or control; randomization balances known and unknown differences.
- Quasi-experimental studies — an intervention compared with a comparison group without random assignment (common in community settings).
- Cohort studies — groups defined by exposure are followed forward in time.
- Case-control studies — people with and without an outcome are compared looking backward.
- Cross-sectional studies — a snapshot of exposure and outcome at one time.
- Case reports, expert opinion — single cases or informed views, useful for hypotheses, not proof.
The hierarchy is a starting point, not a verdict: a poorly done RCT can be less trustworthy than a well-done quasi-experiment. Public health often relies on quasi-experimental and natural experiments because randomizing whole communities is often impossible or unethical — appraise conduct, not label.
Internal validity: is the finding true?
Internal validity asks whether the study's results reflect reality for the people studied. Key questions:
- Was there a comparison? Without one, improvement cannot be attributed to the intervention — the community may have improved anyway.
- Were groups comparable? Randomization helps; otherwise check that groups were similar on age, income, health, and other factors that could explain the result.
- Was follow-up complete? High dropout makes results unrepresentative.
- Were outcomes measured the same way for everyone? Different measurement across groups (information bias) distorts results.
- Could confounding explain it? A third factor tied to both exposure and outcome may be the real driver.
Bias and confounding
- Selection bias: who enters or stays in the study differs systematically between groups (e.g., healthier, more motivated people join the program).
- Information (measurement) bias: outcomes or exposures are recorded inaccurately or differently between groups.
- Confounding: a third variable is tangled with the exposure and independently affects the outcome. People who join an exercise program may also eat better — diet confounds the evaluation of exercise alone. Randomization is the strongest guard; statistical adjustment is only a partial fix.
External validity: will it work here?
A valid study may still not apply to your community. Ask:
- Are the people similar? Age, language, culture, income, and baseline health matter.
- Is the setting similar? Urban versus rural, school versus clinic, high-resource versus low-resource.
- Can we deliver the same dose and quality? Staffing, training, and fidelity affect results.
- Is it feasible and acceptable here? The community may not want or be able to sustain the program.
Evidence transfers with judgment, not automatically.
Effect size, precision, and significance
- Effect size is how big the difference is — a program can be statistically detectable yet trivially small in practice.
- Precision is shown by confidence intervals: a wide interval means the true effect is poorly pinned down; a narrow one means the estimate is stable.
- Statistical significance says the result is unlikely to be due to chance; public health significance asks whether the effect is big enough, in enough people, to matter. Both are needed.
Appraisal tools and reporting standards
You need not reinvent appraisal — checklists and standards exist:
- CASP checklists — free, question-by-question appraisal guides for each study type.
- GRADE System for rating certainty of evidence and recommendation strength Full entry → — a system for rating the certainty of a body of evidence and the strength of guideline recommendations.
- Reporting standards — CONSORT (trials), STROBE (observational studies), PRISMA (reviews) — make it easy to see what a study did; one that omits key details cannot be fully trusted.
- Conflicts of interest: check funding and author disclosures — conflicts can bias even otherwise careful studies.
Common Confusions
| Do Not Confuse | With | Difference |
|---|---|---|
| Internal validity | External validity | Internal asks "is the finding true?"; external asks "will it work here?" — both must pass |
| Statistical significance | Public health significance | Significance says the result is unlikely due to chance; practical significance asks whether the effect is big enough to matter |
| Bias | Confounding | Bias is systematic error in how a study is conducted/measured; confounding is a real third variable tangled with the exposure and outcome |
| Correlation | Causation | Two things occurring together does not prove one causes the other — confounding may explain the link |
| A hierarchy label | Study quality | A poorly run RCT can be worse than a well-run quasi-experiment — appraise conduct, not just design |
| Peer review | Guarantee of truth | Peer review filters obvious errors; flawed papers still pass, and conflicts go undisclosed |

Eli explains
The same idea, in plain words
Explain it like I’m 10
Checking evidence quality is like vetting a recipe before cooking for a crowd. Did the cook test it many times with lots of people (big studies), compare it with a plain version (a comparison group), and honestly list the ingredients (full reporting)? A recipe from one person who never compared anything is risky, however confident they sound.
Worked example
A county health department is considering a community walking program to increase physical activity. Two studies are on the table.
Study A: A local agency evaluated its own walking program and reported that participants walked more after six months — but there was no comparison group, only 40 participants, 30% dropped out, and the vendor funded the report. Without a comparison, improvement can't be attributed to the program; the small sample and dropout threaten representativeness; funding raises a conflict flag. Internal validity fails — suggestive at best.
Study B: A systematic review of 15 trials and quasi-experiments in similar mid-sized communities found walking programs with group-based start-up and built-environment improvements increased walking minutes, with a modest but consistent effect and narrow confidence intervals. One included study was weak; the reviewers noted it and reran analyses without it — results held. Design strong, synthesis transparent (PRISMA), effect consistent and precise: internal validity passes. The county checks external validity — similar demographics, a park system where improvements could be implemented, and residents surveyed in the health assessment support it.
The recommendation: adopt Study B's approach with adaptation and a local evaluation plan — and document why Study A was set aside.
Key takeaways
- Appraise in order: internal validity first, external validity second, effect size third.
- Evidence hierarchy (intervention questions): systematic reviews/meta-analyses → RCTs → quasi-experimental → cohort → case-control → cross-sectional → case reports/expert opinion.
- Randomization is the strongest defense against confounding; quasi-experiments and natural experiments are common in community settings when randomization isn't feasible.
- The big threats: selection bias, information bias, confounding.
- A wide confidence interval = an imprecise estimate; a small effect can still be statistically significant.
- Statistical significance ≠ public health significance — ask "does the size and reach of this effect matter here?"
- Use CASP, GRADE, and reporting standards (CONSORT/STROBE/PRISMA), not gut feeling.
- Check funding and conflicts; flag weak or conflicting evidence for source/SME review rather than overstating it.
Check yourself
5 review questions from the chapter. Try each one, then open the answer.
What are the two questions of appraisal, and in what order are they asked?
Show answer
Internal validity ("is the finding true for the people studied?") first, then external validity ("will it work in our community?").
Rank from strongest to weakest: case-control study, systematic review, cross-sectional study, RCT, Quasi-experimental study Comparison without randomization Full entry →.
Show answer
Systematic review → RCT → quasi-experimental → case-control → cross-sectional.
What is confounding? Give an example relevant to a community health program.
Show answer
Confounding is a third variable associated with both the exposure and the outcome that can explain an apparent effect. Example: people who join a community exercise program may also eat healthier — diet confounds attributing health changes to exercise alone.
Why can't a wide Confidence interval Range that likely contains the true effect Full entry → be ignored even if the result is "statistically significant"?
Show answer
A wide confidence interval means the true effect could be very small or even in the opposite direction — the estimate is imprecise, so acting on it is risky even though the "significance" label is present.
What is the difference between internal and external validity, using a community program as an example?
Show answer
Internal validity: is the study of the walking program conducted well enough that its reported effect is real? External validity: do the study's participants, setting, and program match our community and our capacity to deliver it?
Study tools & related lessonsKey vocabulary · Related
Key vocabulary
- Critical appraisal
- Systematically judging whether a study is valid and applicable
- Internal validity
- Whether the study's result is true for the people studied
- External validity
- Whether the result transfers to other populations and settings
- Bias
- Systematic error that distorts results (selection, information)
- Confounding
- A third variable tangled with exposure and outcome
- Randomized controlled trial
- Experiment with random assignment to intervention/control
- Quasi-experimental study
- Comparison without randomization
- Confidence interval
- Range that likely contains the true effect
- GRADE
- System for rating certainty of evidence and recommendation strength
- Reporting standard
- Checklist for transparent reporting (CONSORT, STROBE, PRISMA)
Sources & references
This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.

