Health Administration · Information

Healthcare Data and Analytics

Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 9 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Quick check
  8. Study tools
  9. Sources & references

In 30 seconds

Healthcare organizations sit on huge amounts of data — clinical records, insurance claims, financial ledgers, and patient-generated readings. Analytics turns that data into decisions. A common ladder runs from descriptive analytics (what happened) to diagnostic (why), predictive (what may happen), and prescriptive (what to do). The insight is only as trustworthy as the data underneath it, and predictive models are regulated decision-support tools, not diagnoses.

Why this matters

Administrators are judged on decisions they can defend with evidence: where to add staff, which patients need outreach, whether a safety problem is real or noise. Analytics is how raw records become those decisions, and knowing its limits keeps you from over-trusting a dashboard. The field is also where health administration, informatics, and clinical care meet, so managers increasingly need to read a risk-stratification report or question a predictive model without being data scientists. As machine learning spreads into care, the ability to ask 'what data trained this, and was it validated?' is becoming a core management skill rather than a technical specialty.

The college version

What counts as healthcare data

Before analytics comes data, and healthcare produces several distinct kinds. describe the patient's health and care: diagnoses, laboratory results, vital signs, medications, imaging, and clinicians' notes. Administrative or claims data are generated for payment — the coded record of what was billed, to whom, for which diagnosis and procedure, on what date, at what price. Financial data track the organization's money: revenue, cost, payroll, and the margins that keep it solvent. come from outside the visit — home blood-pressure cuffs, glucose monitors, wearables, and patient surveys. These sources overlap but are not interchangeable. Claims data cover almost every insured encounter and are cheap to obtain, but they record what was billed rather than what happened clinically, so a diagnosis may be present for reimbursement reasons. Clinical data are richer and closer to the truth of care but are messier and harder to pull together. A capable analytics program knows which source answers which question, and does not treat a billing record as a clinical fact.

Structured vs unstructured data, and where it comes from

Data also differ in form. are standardized and coded — they live in labeled fields a computer can sort and count directly, such as ICD-10-CM diagnosis codes, laboratory values, vital signs, and medication lists. To make records interoperable, ONC maintains the United States Core Data for Interoperability (USCDI), a standardized set of data classes and elements (demographics, laboratory, medications, clinical notes, and more) that certified systems must exchange; USCDI grew from 53 elements in its 2020 first version to over 170 by 2026. are free text and images — the narrative of a progress note, a radiology report, a discharge summary — rich in meaning but hard to analyze without extra processing. Beyond the EHR, standard data sources each have a role: insurance claims for utilization and cost; disease and patient registries that pool cases of a condition; vital statistics from the National Vital Statistics System, which collects official birth and death records — 57 registration jurisdictions send NCHS information on more than 6 million vital events a year; and population surveys such as NHANES and BRFSS that sample the public directly. AHRQ's Quality Indicators, computed from ordinary hospital discharge (claims) data, show how administrative data get reused far beyond billing.

The four analytics tiers

Analytics is usually described as a ladder of four types that answer progressively harder questions. Descriptive analytics reviews historical data to identify patterns — 'what happened' — such as a chart of monthly readmissions or a count of no-shows by clinic. Diagnostic analytics asks 'why did it happen,' digging beneath the summary: does the readmission rate rise with a particular diagnosis, unit, or discharge day? Predictive analytics uses historical data and statistical modeling to forecast future outcomes — 'what may happen' — for example flagging which patients are likely to be readmitted or projecting next week's bed demand. Prescriptive analytics goes furthest, building on diagnostic and predictive results to recommend specific actions, as a clinical decision-support system does when it suggests an intervention. Each rung depends on the ones below it: you cannot usefully predict what you have not first described and understood, and a prescription built on a shaky prediction inherits its errors.

Population health, risk stratification, dashboards, and KPIs

Managers rarely act on one patient at a time; they act on populations. Population health analytics looks across a defined group — a clinic's panel, a health plan's members — to find where need and resources are mismatched. The central technique is : segmenting the population into tiers of similar complexity, such as high-risk, rising-risk, and low-risk, so that scarce care-management effort is aimed at the people most likely to benefit. This matters because a small share of patients typically accounts for a large share of utilization and spending; identifying that group early lets an organization intervene before a crisis rather than after. The results usually surface on a dashboard — a compact display of key performance indicators (KPIs), the handful of metrics leadership watches, like 30-day readmission rate, average length of stay, or emergency-department wait time. A dashboard is descriptive analytics made routine: it tells you the current state at a glance, but it does not by itself explain causes or prescribe action.

Data quality, model bias, and regulated clinical AI

Every layer above rests on — completeness, accuracy, and timeliness. Missing fields, miscoded diagnoses, duplicate records, and delayed feeds all degrade an analysis, and no model can rescue bad inputs; 'garbage in, garbage out' is the field's oldest rule. Predictive models add a second hazard. A machine-learning model learns from historical data, so if that history reflects unequal care, the model can reproduce and even amplify the disparity — that produces unfair outcomes for groups underrepresented or mistreated in the training data. This is why validation matters: a model should be tested on data it did not learn from, and monitored after deployment, before anyone trusts it. Two cautions follow for administrators. First, a predictive score is decision support — it informs a clinician's judgment; it is not a diagnosis and does not act on its own. Second, clinical AI is regulated: the FDA authorizes AI-enabled medical devices for marketing only after a focused review of their safety and effectiveness, maintains a public list of those authorized, and to date radiology accounts for the largest share. Treating a model as a black box that 'just knows' misunderstands both the science and the law.

Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

A hospital collects a mountain of information: what happened to patients, what the insurance bill said, how much money came in, and readings from things like fitness trackers. That pile of facts is data. Analytics is asking the pile good questions. The easiest question is 'What happened?' — like counting how many patients came back within a month. A harder one is 'Why?' Harder still is 'What will happen next?', which needs a prediction. The hardest is 'So what should we do about it?' Managers use this to spot the patients who need the most help and send a nurse before things get worse. But two warnings ride along: if the facts you started with were wrong or missing, every answer is wrong too, and a computer's prediction is a helpful hint for a doctor to weigh, never the final word.

Picture it like this

Think of a weather forecast. Descriptive analytics is the thermometer reading right now; diagnostic is figuring out why it got cold; predictive is the forecast for tomorrow; prescriptive is 'bring an umbrella.' A forecast that helps you plan is not the same as controlling the sky, and it is only as good as the sensors feeding it.

Where the picture stops working

Weather sensors measure physics that does not care who is watching; healthcare data are created by people and billing systems, so they carry human choices and inequities a thermometer never would — a biased model can be unfair in ways a wrong forecast simply is not. And unlike a weather app, clinical prediction tools are regulated medical devices.

Worked example

A hospital wants to understand heart-failure readmissions. Descriptive step: of 1,200 heart-failure discharges last quarter, 168 patients returned within 30 days, so the readmission rate is 168 / 1,200 = 14.0%. Diagnostic step: split by unit — Unit A ran 90 readmissions in 600 discharges (15.0%) while Unit B ran 78 in 600 (13.0%), a ratio of 15.0 / 13.0 ≈ 1.15, meaning Unit A's rate is about 15% higher, worth investigating. Population step: risk-stratify the clinic's panel of 2,500 patients into high-risk (5% = 125), rising-risk (20% = 500), and low-risk (75% = 1,875). With each care manager able to hold about 100 patients, the 125 high-risk patients need roughly two care managers assigned first. Notice every number here is descriptive or diagnostic; deciding to add outreach is prescriptive, and any predictive model that flagged those 125 patients would still need validation before the clinic trusted it.

Key takeaway

Analytics turns healthcare data into decisions along a ladder — describe, diagnose, predict, prescribe — but each rung is only as trustworthy as the data beneath it, and predictive models are regulated, biasable decision-support tools, not diagnoses.

Quick check

3 questions here, of 5 in this lesson’s practice set. Answers stay hidden until you check.

Question 1 of 3intermediate

A hospital reports that 168 of 1,200 heart-failure patients were readmitted within 30 days last quarter and displays the 14% rate on its leadership dashboard. Which analytics tier does this display represent?

Choose an answer, then check it.
Question 2 of 3intermediate

An analyst needs data on almost every insured encounter — diagnoses, procedures, dates, and charges — to study utilization and cost across a health plan. Which data source is the natural first choice, and what is its key limitation?

Choose an answer, then check it.
Question 3 of 3foundational

Which pairing correctly separates structured from unstructured healthcare data?

Choose an answer, then check it.
Practice all 5

Keep learning

Ready to build on this? Continue to the next lesson.

Practice this lesson
Study tools & related lessonsYou’ll learn to · Common mistakes · Easily confused · Key vocabulary · Related

You’ll learn to

  • Distinguish the main types of healthcare data: clinical, administrative/claims, financial, and patient-generated.
  • Explain the four analytics tiers — descriptive, diagnostic, predictive, prescriptive — with healthcare examples.
  • Identify key data sources (claims, EHR data, registries, vital statistics, surveys) and tell structured from unstructured data.
  • Apply risk stratification to segment a population and interpret a simple rate or ratio.
  • Evaluate how data quality and model bias limit what analytics can support, and describe how clinical AI is regulated.

Common mistakes

  • Treating claims data as clinical truth.

    Claims are created to get paid. A code may be present, absent, or shaped by reimbursement rules, so a claim tells you what was billed, not exactly what happened at the bedside.

  • Calling a dashboard 'predictive analytics.'

    A dashboard of current KPIs is descriptive analytics — it reports the present state. Prediction requires a model that forecasts a future outcome, which most dashboards do not do.

  • Assuming more data automatically means better answers.

    Analytics is bounded by data quality. Incomplete, miscoded, duplicated, or stale data produce confident but wrong conclusions no matter how large the dataset.

  • Trusting a predictive model as if it were objective or a diagnosis.

    Models learn from historical data and can inherit its biases; a risk score is decision support that informs a clinician, must be validated on new data, and — for clinical AI — is regulated by the FDA.

  • Confusing risk stratification with rationing.

    Stratification matches level of care to level of need so limited resources reach the people most likely to benefit; it guides intensity of support, not denial of care.

Easily confused

Claims/administrative data vs. Clinical/EHR data

Claims are billing-driven, broad, and cheap but shallow; clinical data are care-driven, deep, and truer but messier and harder to aggregate.

Structured data vs. Unstructured data

Structured data are coded fields a computer counts directly; unstructured data are free text and images that need processing before analysis.

Predictive analytics vs. Prescriptive analytics

Prediction forecasts what may happen; prescription recommends what to do about it and depends on the prediction being sound.

Descriptive dashboard/KPI vs. Predictive model

A dashboard reports the current state; a model estimates a future outcome and, in clinical use, is a regulated, validated decision-support tool.

Key vocabulary

Clinical data
Information about a patient's health and care — diagnoses, lab results, vital signs, medications, imaging, and notes — generated during the delivery of care.
Claims (administrative) data
The coded record created for billing and payment, capturing diagnoses, procedures, dates, sites, and charges for an encounter rather than the full clinical story.
Patient-generated data
Health data produced outside the clinical encounter, such as readings from wearables and home monitors or answers to patient surveys.
Structured data
Standardized, coded values stored in labeled fields (e.g., ICD-10 codes, lab values, vital signs) that software can sort and count directly.
Unstructured data
Free-text and image content such as clinical notes, radiology reports, and discharge summaries, which requires extra processing before it can be analyzed.
Descriptive / diagnostic / predictive / prescriptive analytics
The four analytics tiers that answer, in order, what happened, why it happened, what may happen, and what action to take.
Risk stratification
Segmenting a population into tiers of similar complexity or risk (e.g., high, rising, low) so care-management resources can be prioritized.
Key performance indicator (KPI)
A metric leadership tracks to gauge performance, such as 30-day readmission rate or average length of stay; KPIs are the content of most dashboards.
Data quality
The degree to which data are complete, accurate, and timely enough to be trusted; poor quality limits every analysis built on it.
Algorithmic bias
Systematic unfairness in a model's outputs that arises when historical training data reflect existing disparities, which the model then reproduces or amplifies.

Sources & references

  1. Healthcare Analytics (StatPearls) — StatPearls / NCBI Bookshelf
  2. Artificial Intelligence-Enabled Medical Devices (AI-Enabled Medical Device List) — U.S. Food and Drug Administration (Center for Devices and Radiological Health)
  3. About the National Vital Statistics System (NVSS) — CDC / National Center for Health Statistics (NCHS)
  4. United States Core Data for Interoperability (USCDI) — ONC / HealthIT.gov (Assistant Secretary for Technology Policy)
  5. AHRQ Quality Indicators — Agency for Healthcare Research and Quality (AHRQ)
  6. Population risk stratification tools and interventions for chronic disease management in primary care: a systematic literature review — PMC (peer-reviewed systematic review)
  7. ICD-10-CM/PCS Transition: Background (National Center for Health Statistics) — CDC / National Center for Health Statistics

EliExplains lessons are original prose written from the open, credible references above. See Copyright & Licensing.

Researched 2026-08-19

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.