Health Administration · Information
Healthcare Data and Analytics
On this page 9 sections
In 30 seconds
Healthcare organizations sit on huge amounts of data — clinical records, insurance claims, financial ledgers, and patient-generated readings. Analytics turns that data into decisions. A common ladder runs from descriptive analytics (what happened) to diagnostic (why), predictive (what may happen), and prescriptive (what to do). The insight is only as trustworthy as the data underneath it, and predictive models are regulated decision-support tools, not diagnoses.
Why this matters
Administrators are judged on decisions they can defend with evidence: where to add staff, which patients need outreach, whether a safety problem is real or noise. Analytics is how raw records become those decisions, and knowing its limits keeps you from over-trusting a dashboard. The field is also where health administration, informatics, and clinical care meet, so managers increasingly need to read a risk-stratification report or question a predictive model without being data scientists. As machine learning spreads into care, the ability to ask 'what data trained this, and was it validated?' is becoming a core management skill rather than a technical specialty.
The college version
What counts as healthcare data
Before analytics comes data, and healthcare produces several distinct kinds. Clinical data Information about a patient's health and care — diagnoses, lab results, vital signs, medications, imaging, and notes — generated during the delivery of care. Full entry → describe the patient's health and care: diagnoses, laboratory results, vital signs, medications, imaging, and clinicians' notes. Administrative or claims data are generated for payment — the coded record of what was billed, to whom, for which diagnosis and procedure, on what date, at what price. Financial data track the organization's money: revenue, cost, payroll, and the margins that keep it solvent. Patient-generated data Health data produced outside the clinical encounter, such as readings from wearables and home monitors or answers to patient surveys. Full entry → come from outside the visit — home blood-pressure cuffs, glucose monitors, wearables, and patient surveys. These sources overlap but are not interchangeable. Claims data cover almost every insured encounter and are cheap to obtain, but they record what was billed rather than what happened clinically, so a diagnosis may be present for reimbursement reasons. Clinical data are richer and closer to the truth of care but are messier and harder to pull together. A capable analytics program knows which source answers which question, and does not treat a billing record as a clinical fact.
Structured vs unstructured data, and where it comes from
Data also differ in form. Structured data Standardized, coded values stored in labeled fields (e.g., ICD-10 codes, lab values, vital signs) that software can sort and count directly. Full entry → are standardized and coded — they live in labeled fields a computer can sort and count directly, such as ICD-10-CM diagnosis codes, laboratory values, vital signs, and medication lists. To make records interoperable, ONC maintains the United States Core Data for Interoperability (USCDI), a standardized set of data classes and elements (demographics, laboratory, medications, clinical notes, and more) that certified systems must exchange; USCDI grew from 53 elements in its 2020 first version to over 170 by 2026. Unstructured data Free-text and image content such as clinical notes, radiology reports, and discharge summaries, which requires extra processing before it can be analyzed. Full entry → are free text and images — the narrative of a progress note, a radiology report, a discharge summary — rich in meaning but hard to analyze without extra processing. Beyond the EHR, standard data sources each have a role: insurance claims for utilization and cost; disease and patient registries that pool cases of a condition; vital statistics from the National Vital Statistics System, which collects official birth and death records — 57 registration jurisdictions send NCHS information on more than 6 million vital events a year; and population surveys such as NHANES and BRFSS that sample the public directly. AHRQ's Quality Indicators, computed from ordinary hospital discharge (claims) data, show how administrative data get reused far beyond billing.
The four analytics tiers
Analytics is usually described as a ladder of four types that answer progressively harder questions. Descriptive analytics reviews historical data to identify patterns — 'what happened' — such as a chart of monthly readmissions or a count of no-shows by clinic. Diagnostic analytics asks 'why did it happen,' digging beneath the summary: does the readmission rate rise with a particular diagnosis, unit, or discharge day? Predictive analytics uses historical data and statistical modeling to forecast future outcomes — 'what may happen' — for example flagging which patients are likely to be readmitted or projecting next week's bed demand. Prescriptive analytics goes furthest, building on diagnostic and predictive results to recommend specific actions, as a clinical decision-support system does when it suggests an intervention. Each rung depends on the ones below it: you cannot usefully predict what you have not first described and understood, and a prescription built on a shaky prediction inherits its errors.
Population health, risk stratification, dashboards, and KPIs
Managers rarely act on one patient at a time; they act on populations. Population health analytics looks across a defined group — a clinic's panel, a health plan's members — to find where need and resources are mismatched. The central technique is Risk stratification Segmenting a population into tiers of similar complexity or risk (e.g., high, rising, low) so care-management resources can be prioritized. Full entry →: segmenting the population into tiers of similar complexity, such as high-risk, rising-risk, and low-risk, so that scarce care-management effort is aimed at the people most likely to benefit. This matters because a small share of patients typically accounts for a large share of utilization and spending; identifying that group early lets an organization intervene before a crisis rather than after. The results usually surface on a dashboard — a compact display of key performance indicators (KPIs), the handful of metrics leadership watches, like 30-day readmission rate, average length of stay, or emergency-department wait time. A dashboard is descriptive analytics made routine: it tells you the current state at a glance, but it does not by itself explain causes or prescribe action.
Data quality, model bias, and regulated clinical AI
Every layer above rests on Data quality The degree to which data are complete, accurate, and timely enough to be trusted; poor quality limits every analysis built on it. Full entry → — completeness, accuracy, and timeliness. Missing fields, miscoded diagnoses, duplicate records, and delayed feeds all degrade an analysis, and no model can rescue bad inputs; 'garbage in, garbage out' is the field's oldest rule. Predictive models add a second hazard. A machine-learning model learns from historical data, so if that history reflects unequal care, the model can reproduce and even amplify the disparity — Algorithmic bias Systematic unfairness in a model's outputs that arises when historical training data reflect existing disparities, which the model then reproduces or amplifies. Full entry → that produces unfair outcomes for groups underrepresented or mistreated in the training data. This is why validation matters: a model should be tested on data it did not learn from, and monitored after deployment, before anyone trusts it. Two cautions follow for administrators. First, a predictive score is decision support — it informs a clinician's judgment; it is not a diagnosis and does not act on its own. Second, clinical AI is regulated: the FDA authorizes AI-enabled medical devices for marketing only after a focused review of their safety and effectiveness, maintains a public list of those authorized, and to date radiology accounts for the largest share. Treating a model as a black box that 'just knows' misunderstands both the science and the law.

Eli explains
The same idea, in plain words
Explain it like I’m 10
A hospital collects a mountain of information: what happened to patients, what the insurance bill said, how much money came in, and readings from things like fitness trackers. That pile of facts is data. Analytics is asking the pile good questions. The easiest question is 'What happened?' — like counting how many patients came back within a month. A harder one is 'Why?' Harder still is 'What will happen next?', which needs a prediction. The hardest is 'So what should we do about it?' Managers use this to spot the patients who need the most help and send a nurse before things get worse. But two warnings ride along: if the facts you started with were wrong or missing, every answer is wrong too, and a computer's prediction is a helpful hint for a doctor to weigh, never the final word.
Picture it like this
Think of a weather forecast. Descriptive analytics is the thermometer reading right now; diagnostic is figuring out why it got cold; predictive is the forecast for tomorrow; prescriptive is 'bring an umbrella.' A forecast that helps you plan is not the same as controlling the sky, and it is only as good as the sensors feeding it.
Where the picture stops working
Weather sensors measure physics that does not care who is watching; healthcare data are created by people and billing systems, so they carry human choices and inequities a thermometer never would — a biased model can be unfair in ways a wrong forecast simply is not. And unlike a weather app, clinical prediction tools are regulated medical devices.
Worked example
A hospital wants to understand heart-failure readmissions. Descriptive step: of 1,200 heart-failure discharges last quarter, 168 patients returned within 30 days, so the readmission rate is 168 / 1,200 = 14.0%. Diagnostic step: split by unit — Unit A ran 90 readmissions in 600 discharges (15.0%) while Unit B ran 78 in 600 (13.0%), a ratio of 15.0 / 13.0 ≈ 1.15, meaning Unit A's rate is about 15% higher, worth investigating. Population step: risk-stratify the clinic's panel of 2,500 patients into high-risk (5% = 125), rising-risk (20% = 500), and low-risk (75% = 1,875). With each care manager able to hold about 100 patients, the 125 high-risk patients need roughly two care managers assigned first. Notice every number here is descriptive or diagnostic; deciding to add outreach is prescriptive, and any predictive model that flagged those 125 patients would still need validation before the clinic trusted it.
Key takeaway
Analytics turns healthcare data into decisions along a ladder — describe, diagnose, predict, prescribe — but each rung is only as trustworthy as the data beneath it, and predictive models are regulated, biasable decision-support tools, not diagnoses.
Quick check
3 questions here, of 5 in this lesson’s practice set. Answers stay hidden until you check.
An analyst needs data on almost every insured encounter — diagnoses, procedures, dates, and charges — to study utilization and cost across a health plan. Which data source is the natural first choice, and what is its key limitation?
Which pairing correctly separates structured from unstructured healthcare data?
Study tools & related lessonsYou’ll learn to · Common mistakes · Easily confused · Key vocabulary · Related
You’ll learn to
- Distinguish the main types of healthcare data: clinical, administrative/claims, financial, and patient-generated.
- Explain the four analytics tiers — descriptive, diagnostic, predictive, prescriptive — with healthcare examples.
- Identify key data sources (claims, EHR data, registries, vital statistics, surveys) and tell structured from unstructured data.
- Apply risk stratification to segment a population and interpret a simple rate or ratio.
- Evaluate how data quality and model bias limit what analytics can support, and describe how clinical AI is regulated.
Common mistakes
Treating claims data as clinical truth.
Claims are created to get paid. A code may be present, absent, or shaped by reimbursement rules, so a claim tells you what was billed, not exactly what happened at the bedside.
Calling a dashboard 'predictive analytics.'
A dashboard of current KPIs is descriptive analytics — it reports the present state. Prediction requires a model that forecasts a future outcome, which most dashboards do not do.
Assuming more data automatically means better answers.
Analytics is bounded by data quality. Incomplete, miscoded, duplicated, or stale data produce confident but wrong conclusions no matter how large the dataset.
Trusting a predictive model as if it were objective or a diagnosis.
Models learn from historical data and can inherit its biases; a risk score is decision support that informs a clinician, must be validated on new data, and — for clinical AI — is regulated by the FDA.
Confusing risk stratification with rationing.
Stratification matches level of care to level of need so limited resources reach the people most likely to benefit; it guides intensity of support, not denial of care.
Easily confused
Claims/administrative data vs. Clinical/EHR data
Claims are billing-driven, broad, and cheap but shallow; clinical data are care-driven, deep, and truer but messier and harder to aggregate.
Structured data vs. Unstructured data
Structured data are coded fields a computer counts directly; unstructured data are free text and images that need processing before analysis.
Predictive analytics vs. Prescriptive analytics
Prediction forecasts what may happen; prescription recommends what to do about it and depends on the prediction being sound.
Descriptive dashboard/KPI vs. Predictive model
A dashboard reports the current state; a model estimates a future outcome and, in clinical use, is a regulated, validated decision-support tool.
Key vocabulary
- Clinical data
- Information about a patient's health and care — diagnoses, lab results, vital signs, medications, imaging, and notes — generated during the delivery of care.
- Claims (administrative) data
- The coded record created for billing and payment, capturing diagnoses, procedures, dates, sites, and charges for an encounter rather than the full clinical story.
- Patient-generated data
- Health data produced outside the clinical encounter, such as readings from wearables and home monitors or answers to patient surveys.
- Structured data
- Standardized, coded values stored in labeled fields (e.g., ICD-10 codes, lab values, vital signs) that software can sort and count directly.
- Unstructured data
- Free-text and image content such as clinical notes, radiology reports, and discharge summaries, which requires extra processing before it can be analyzed.
- Descriptive / diagnostic / predictive / prescriptive analytics
- The four analytics tiers that answer, in order, what happened, why it happened, what may happen, and what action to take.
- Risk stratification
- Segmenting a population into tiers of similar complexity or risk (e.g., high, rising, low) so care-management resources can be prioritized.
- Key performance indicator (KPI)
- A metric leadership tracks to gauge performance, such as 30-day readmission rate or average length of stay; KPIs are the content of most dashboards.
- Data quality
- The degree to which data are complete, accurate, and timely enough to be trusted; poor quality limits every analysis built on it.
- Algorithmic bias
- Systematic unfairness in a model's outputs that arises when historical training data reflect existing disparities, which the model then reproduces or amplifies.
Sources & references
- Healthcare Analytics (StatPearls) — StatPearls / NCBI Bookshelf
- Artificial Intelligence-Enabled Medical Devices (AI-Enabled Medical Device List) — U.S. Food and Drug Administration (Center for Devices and Radiological Health)
- About the National Vital Statistics System (NVSS) — CDC / National Center for Health Statistics (NCHS)
- United States Core Data for Interoperability (USCDI) — ONC / HealthIT.gov (Assistant Secretary for Technology Policy)
- AHRQ Quality Indicators — Agency for Healthcare Research and Quality (AHRQ)
- Population risk stratification tools and interventions for chronic disease management in primary care: a systematic literature review — PMC (peer-reviewed systematic review)
- ICD-10-CM/PCS Transition: Background (National Center for Health Statistics) — CDC / National Center for Health Statistics
EliExplains lessons are original prose written from the open, credible references above. See Copyright & Licensing.
Researched 2026-08-19
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.

