Data Science & AI Literacy · Foundations

Large Language Models

Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 9 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Quick check
  8. Study tools
  9. Sources & references

In 30 seconds

A , or LLM, is a deep learning model trained on enormous amounts of text — books, web pages, and code — so it can understand and generate natural language. It learns patterns by predicting what comes next, and generating text is just that choice repeated, one at a time. As of 2026 these models power chatbots that summarize, draft, translate, explain, and assist with code. Their fluent output is not guaranteed fact, so it deserves checking.

Why this matters

Large language models are the technology behind the chatbots and writing assistants people meet daily, so knowing what they are and how they work changes how you read their output. Understanding that an LLM predicts patterns rather than consulting facts explains both why it can write fluently and why it can be confidently wrong. That distinction matters in school, at work, and for anyone evaluating information: it tells you when an answer is worth acting on and when it needs verification. It also gives you language for talking about AI critically rather than treating it as magic.

The college version

What a large language model is

A large language model, or LLM, is a deep learning model trained on immense amounts of text that learns to understand and generate natural language. IBM's public explainer, read in August 2026, describes LLMs as deep learning models trained on immense amounts of data, capable of understanding and generating natural language, built on a architecture that excels at handling sequences of words. OpenAI's documentation makes the practical point: you use a large language model to generate text from a , as you might with ChatGPT. The name carries the definition: 'language' says the material is text, 'model' says it is a trained system rather than a database or hand-written rules, and 'large' points at the defining ingredient — training on far more text than any person reads in a lifetime. The assistants sold as Claude, ChatGPT, Copilot, Llama, and Gemini, among the interfaces IBM names, are all built on LLMs. Everything a model can do comes from the patterns it absorbed during training.

How it works: prediction, one token at a time

The core mechanism is simple: predict what comes next. IBM describes LLMs as giant statistical prediction machines that repeatedly predict the next word in a sequence, having learned patterns in text so they generate language that follows those patterns. Text is first broken into small units called tokens — words, subwords, or characters. During training, the model practices one task over and over: given a stretch of text, guess the token that follows. Each wrong guess nudges its internal settings, until the model has captured the patterns of grammar, word choice, and structure in its training material. Generating works the same way. Given 'The cat sat on the,' the model assigns likelihoods to candidate next tokens and picks one, then treats its own choice as context and picks the token after that, repeating until it stops. Text generation is just a long chain of choosing the next token, informed by everything before it. No step looks anything up, and no step consults a fact-checker.

Training at scale: books, web pages, and code

Scale is what separates an LLM from earlier language programs. IBM describes training as starting with a massive amount of data — billions or trillions of words from books, articles, websites, code, and other text sources — cleaned before use. The transformer architecture, introduced in 2017, made training on datasets this large practical, and the resulting models hold billions or trillions of internal settings, or parameters. Because the collection is so broad, the model picks up patterns far beyond grammar: how explanations are structured, what a summary looks like, how instructions are followed, what code tends to look like. That is why the same model can summarize an article and help debug a program. 'Large' is not decoration — the size and variety of the training text is the defining feature of the approach.

What LLMs do well, as of 2026

As of the sources read in August 2026, LLMs handle a broad set of everyday language tasks. IBM lists interpreting and generating text for jobs like summarizing an article, debugging code, and drafting a legal clause, plus translation and multilingual capabilities and explaining complex concepts in simpler terms. OpenAI's documentation observes that models can generate almost any kind of text response — code, mathematical equations, structured data, or human-like prose. Examples of the general capability: Priya pastes a dense forty-page journal article and asks for a three-paragraph plain-language summary; Marco asks a coding assistant why his script crashes on empty input; a shopkeeper asks for a friendlier refund email; a student asks for a hard chapter explained to a ten-year-old. Tools and quality keep changing, but the general capability — turning a request into fluent, useful-sounding text — is what the providers document as of 2026.

What LLMs get wrong

Fluent output is not the same as true output. IBM lists accuracy as a major concern: during hallucinations the model generates information that is false or misleading while sounding plausible. NIST's Generative AI Profile documents the same failure mode — confabulation — plausible-sounding text that is simply wrong, known colloquially as hallucinations — as a risk that can mislead users. The mechanism explains why: a plausible-sounding citation fits the pattern of the question beautifully, whether or not the cited work exists, so a chatbot can invent a study, an author, or a statistic with complete confidence. The hallucinations sibling topic covers this in depth. The practical boundary: an LLM is a pattern-continuation engine, and patterns are not proof. IBM also notes that models can reflect and amplify biases in their training data — the corpus shapes the model, for good and for ill.

LLMs, generative AI, and using them well

Large language models sit inside the broader family of generative AI — AI that creates new content rather than only classifying or predicting. The boundary is one sentence: generative AI covers text, images, audio, video, and code, while LLMs are the text-focused family within it, which IBM describes as the most common kind of foundation model, built for text generation. Microsoft Learn's generative AI module likewise treats LLMs as a core generative AI concept. Using an LLM well comes down to two habits covered by sibling topics: write effective prompts — what you ask and how you frame it changes what you get (prompting-basics) — and verify output, checking facts, citations, and numbers against reliable sources because confabulation is documented (evaluating-ai-output). An LLM is a fast, fluent writer with no fact-checker built in; the human at the keyboard supplies the checking.

Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

A large language model is a program that learned to talk by reading an enormous amount of writing — books, web pages, and code — and practicing one simple trick: guess what comes next. Given 'The dog chased the,' it guesses 'ball' is likely, then guesses what follows 'ball,' and keeps going until it has produced a whole answer. Every sentence a chatbot writes is built that way, one small piece at a time. Because it practiced on so much text, it got very good at writing things that sound right. But sounding right is not the same as being right, and the model never checks its work, so it can confidently invent things that do not exist.

Picture it like this

Imagine a stage actor who has read thousands of books and plays, and whose party trick is finishing any story you start. You say, 'A sailor walks into a café and orders...' and he continues fluently, in character, without pausing. He is not remembering one particular book — he is drawing on everything he has read to choose what would most naturally come next, again and again, until the story lands somewhere. An LLM does exactly that with text: it has 'read' a vast library and its skill is continuing any passage plausibly.

Where the picture stops working

The actor knows the difference between a story and a fact: if you ask whether the café really exists, he will say no. A language model cannot make that distinction — it continues patterns whether the topic is fiction or physics, so it will invent a 'fact' as readily as a story. The actor also understands what he says; the model only chooses what fits. And while the actor knows when he is improvising, the model has no such awareness, which is why its confident mistakes need checking.

Worked example

Devon is writing a class report on recycling programs and wants a plain-language summary of a dense city report. He pastes the text into a chatbot and asks for a five-sentence summary; the bot breaks the request into tokens and produces a fluent paragraph — generation by repeated next-token choices. The summary is accurate where Devon checks it against the original, so he keeps it. Then he asks the bot to add citations, and it supplies 'a 2022 study by the Greenfield Institute' with a journal name and a page number. Devon searches for it and finds nothing — the citation is a hallucination, a confident pattern that matches the shape of a real reference. He deletes the fabricated citation, verifies the remaining claims against the city report itself, and submits his paper with only sources he can actually find.

Key takeaway

A large language model is a text-prediction machine trained on enormous amounts of text: it generates fluent language by choosing the next token, again and again — and because it continues patterns rather than checking facts, its confident output needs human verification.

Quick check

3 questions here, of 5 in this lesson’s practice set. Answers stay hidden until you check.

Question 1 of 3foundational

Which of the following best describes a large language model?

Choose an answer, then check it.
Question 2 of 3intermediate

A language model is asked to finish the sentence 'The mail carrier left the package at the...' and produces 'door.' Which best describes the mechanism behind that word?

Choose an answer, then check it.
Question 3 of 3intermediate

A student asks a chatbot to explain a difficult physics chapter 'as if to a ten-year-old,' and the chatbot produces a clear, simple explanation. As of the 2026 sources, what best explains why the model can do this?

Choose an answer, then check it.
Practice all 5

Keep learning

Ready to build on this? Continue to the next lesson.

Practice this lesson
Study tools & related lessonsYou’ll learn to · Common mistakes · Easily confused · Key vocabulary · Related

You’ll learn to

  • Define a large language model as a model trained on vast amounts of text to understand and generate natural language.
  • Explain the generation mechanism: an LLM predicts the next token from its context, and generating text is repeating that choice.
  • Describe what makes these models 'large': training on enormous text corpora — books, web pages, and code — is the defining feature.
  • Give time-anchored examples of tasks LLMs handled well as of the 2026 sources: summarizing, drafting, translating, explaining, and assisting with code.
  • Identify the documented failure mode of confident but wrong output, including fabricated citations, and explain why fluency is not accuracy.
  • Distinguish LLMs as the text-focused family of generative AI and apply a basic verification habit to their output.

Common mistakes

  • Treating a chatbot like a search engine that looks up answers.

    A search engine returns pages that already exist; an LLM generates new text by predicting likely next tokens from learned patterns. Fluency is not retrieval, which is why a confident answer is not proof the underlying fact exists.

  • Believing that a confident, well-written answer is accurate.

    Fluency is not accuracy. IBM documents hallucinations — false or misleading output that sounds plausible — as a major concern, and NIST lists confabulation as a documented risk. Citations and numbers deserve checking against real sources.

  • Thinking 'large' refers to how much the model has memorized.

    The 'large' in large language model points to training at scale — billions or trillions of words from books, web pages, and code. The model learns patterns from that corpus; it does not memorize the corpus like a database.

  • Assuming every AI chatbot is the same kind of system.

    Chatbots are interfaces; the underlying systems differ. An LLM generates text from patterns, while other AI systems classify, predict, or retrieve. The label 'AI' alone does not tell you which kind you are using.

  • Using an LLM's answer without verifying the important parts.

    Because confabulation is documented, facts, citations, and numbers deserve a second look against reliable sources. The evaluating-ai-output topic covers how to check AI output systematically.

Easily confused

Large language model vs. Search engine

A search engine retrieves existing pages that match your query; an LLM generates new text by predicting likely next tokens from patterns learned during training. One finds what exists; the other produces what sounds right.

Large language model vs. Other generative AI (images, audio, video)

All are generative AI that create new content from learned patterns. LLMs are the text-focused family — IBM calls them the most common foundation models, built for text generation — while image, audio, and video models generate those other kinds of content.

Large language model vs. Rule-based language program

A rule-based program follows grammar rules written by people, so it breaks on anything the rules did not anticipate. An LLM learned its patterns by predicting tokens across billions of words of text, which is why it handles messy, varied language — but it has no rulebook to keep it honest.

Key vocabulary

large language model
A deep learning model trained on vast amounts of text that learns to understand and generate natural language, typically by predicting what comes next.
token
A small unit of text — a word, subword, or character — that a language model reads and generates one at a time.
next-token prediction
The training task of guessing the next token in a sequence from the tokens that come before it, repeated millions of times during training.
training corpus
The enormous collection of text — books, articles, websites, and code — that a language model learns its patterns from.
transformer
The neural network architecture, introduced in 2017, that large language models are built on and that made training on very large text datasets practical.
parameter
An internal setting of a model that is adjusted during training and controls how the model processes text; large models hold billions or trillions of them.
prompt
The request a user gives a language model to tell it what text to produce; prompting-basics covers how to write prompts well.
hallucination
A confidently stated but false output, such as an invented citation, produced when a model continues patterns instead of checking facts; the hallucinations topic covers this in depth.

Sources & references

  1. What are large language models (LLMs)? — IBM
  2. Text generation (OpenAI API documentation) — OpenAI
  3. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1) — National Institute of Standards and Technology (NIST)
  4. Introduction to generative AI and agents (Microsoft Learn training module: fundamentals-generative-ai) — Microsoft Learn

EliExplains lessons are original prose written from the open, credible references above. See Copyright & Licensing.

Researched 2026-08-21

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.