Computer Literacy · Foundations
Search Engines
On this page 9 sections
In 30 seconds
A web Search engine Software that finds web pages matching a query by searching its own pre-built index of the web and returning ranked results. Full entry → is software that finds pages for a query in three stages: crawling, where automated crawlers discover and download pages by following links; indexing, where it analyzes those pages and stores them in a huge database; and serving, where it ranks the indexed pages by relevance and quality for your search. You get better results with specific keywords, quotes for exact phrases, and operators like site:. On the results page, learn to tell labeled ads from organic results and to judge each source.
Why this matters
Search engines are how most people reach information, so knowing how they work changes what you do with them. If you understand crawling, indexing, and Ranking The programmatic ordering of indexed pages for a specific query, based on estimated relevance and quality across many factors. Full entry →, you know why a page you saw yesterday is gone, why the top result is not automatically the best answer, and why a labeled ad sits above the organic results. Good operators turn a vague search into a precise one, saving time in coursework and research. Understanding that results are ranked, optimized by site owners, and personalized to your location and language helps you read them critically instead of trusting the order. These habits carry into every class, job, and everyday decision that starts with a search box.
The college version
The three-stage pipeline: crawl, index, serve
A general web search engine does not search the live internet when you type a query. It searches its own stored copy of the web, built ahead of time in three stages. The first is crawling. Automated programs called crawlers (Google's is Googlebot) download the text, images, and video from pages they find online. They discover new pages mainly by following links from pages they already know, by reading sitemaps that site owners submit, and by revisiting pages to check for changes. An algorithmic process decides which sites to crawl, how often, and how many pages to fetch, so the crawler cannot and does not read the entire web at once. The second stage is indexing. The engine analyzes each crawled page, works out what it is about, and stores that information in an Index The large database in which a search engine stores information about the pages it has analyzed, so they can be looked up quickly for a query. Full entry →, a very large database organized so pages can be looked up quickly by their words and features. Being crawled does not guarantee being indexed: not every page an engine processes is added to the index, and problems such as blocked access or thin, low-quality content can keep a page out. The third stage is serving. When you search, the engine looks in its index, not the live web, and returns the pages it judges most useful, ranked in order. This is why an engine can answer in a fraction of a second, and why a page can appear in results even after the original site goes down: you may be seeing what the index stored.
Searching well: keywords, phrases, and operators
The quality of a search depends heavily on the query. Specific, distinctive keywords beat vague ones, because they match fewer, more relevant pages. A few standard refinements sharpen results further. Putting a phrase inside quotation marks, like "first past the post", tells the engine to match that exact wording rather than the words scattered separately. Placing a minus sign directly before a word excludes it: jaguar -car pushes away the automobile and toward the animal. The site: operator limits results to one website or domain, so site:nih.gov diabetes searches only that domain. The filetype: operator finds a particular document format, so filetype:pdf returns PDF files. Many engines also support date filters such as before: and after:. You do not need to memorize a long list; knowing that these tools exist, and reaching for quotes, minus, and site: when a search is too broad or too noisy, is what separates a quick, precise search from endless scrolling. These operators are widely supported, though the exact set varies by engine.
Reading the results page: ads, organic results, and credibility
A results page is not one neutral list. Some entries are advertisements: a business paid for placement, usually at the top or bottom, and these carry a label such as "Sponsored" or "Ad." The rest are organic (or natural) results, ranked by the engine's own judgment of relevance and quality rather than payment. The distinction matters enough that the U.S. Federal Trade Commission has told search engines that paid results must be clearly distinguishable from natural results, and that failing to clearly and prominently distinguish advertising could be a deceptive practice. So the first habit on a results page is to notice the ad labels and know that a paid position is not a vote of quality. The second habit is to evaluate the actual source you click. A common classroom framework is the CRAAP test, which asks five questions about a source: Currency (is it up to date?), Relevance (does it fit your need?), Authority (is the author or organization qualified?), Accuracy (is it supported and verifiable?), and Purpose (why was it made: to inform, to persuade, or to sell?). Ranking answers popularity and relevance, not truth, so the reader still has to judge the source.
Results are ranked, optimized, and personalized
Three facts explain why the order of results is not a simple measure of what is best. First, results are ranked programmatically by many factors that estimate relevance and quality; Google states it does not accept payment to rank pages higher in its organic results, so organic order reflects the algorithm, not a purchase. Second, site owners practice Search engine optimization (SEO) The practice of adjusting a site to help search engines understand its content and users find it; it improves the chance of ranking but guarantees nothing. Full entry →, which Google describes as helping search engines understand a site's content and helping users find it. Legitimate SEO improves clarity and quality, but it also means many results you see were shaped to rank well, and no technique guarantees a top position or even that a page will be indexed. Third, results are personalized and contextual: ranking factors can include your location, language, and device, so two people running the same query can see different orderings. There is no single neutral list handed to everyone. Reading results well means holding all three in mind: the order is an algorithm's estimate, influenced by optimization, and tailored to you.

Eli explains
The same idea, in plain words
Explain it like I’m 10
A search engine does not run around the whole internet the moment you ask a question. Long before you search, it sends out little programs that follow links from page to page and copy down what they find. It reads all those copies and files them in a giant catalog. When you type your question, it flips through its own catalog, not the live internet, and hands you the pages it thinks fit best, in order. That is why answers come back so fast. The list you get mixes paid ads, which are labeled, with regular results the engine chose. The top spot means the engine guessed it is a good match, not that it is true or the only answer, so you still check who wrote each page.
Picture it like this
It works like a huge library with a card catalog. Helpers (crawlers) walk the shelves and write a card for every book (indexing). When you ask a question at the desk, the librarian searches the cards, not every shelf, and points you to the best-matching books (ranking). A few books on display up front are there because a publisher paid for the spot, not because they are best.
Where the picture stops working
The library breaks down in three ways: the web changes constantly and pages vanish or move, so the catalog is always a bit out of date; the engine tailors its answer to you, while a real librarian gives everyone the same shelf; and books are vetted before shelving, whereas anyone can publish a web page, so you still have to judge each source yourself.
Worked example
Suppose you need a recent, official statistic on U.S. high school graduation rates for a paper, and a plain search for graduation rates buries you in coaching sites and ads. Refine it in steps. Add the exact phrase in quotes: "graduation rate". Limit the domain to a government source with the site: operator: site:nces.ed.gov. Ask for a report file with filetype:pdf. Your query becomes "graduation rate" site:nces.ed.gov filetype:pdf. Now scan the results: skip anything labeled Sponsored, since that is a paid slot, and open an organic result from the agency. Before quoting it, run a quick CRAAP check, is it current, authoritative, and meant to inform rather than sell? A government statistics agency page passes, so you cite it.
Key takeaway
A search engine crawls the web, indexes what it finds, and ranks that index for your query, so the results are an algorithm's tailored estimate, not a neutral list of truth. Search precisely with operators, tell ads from organic results, and judge each source.
Quick check
3 questions here, of 5 in this lesson’s practice set. Answers stay hidden until you check.
During the crawling stage, what does a search engine's crawler mainly do?
You want results only from a single website, nih.gov. Which refinement does that directly?
Study tools & related lessonsYou’ll learn to · Common mistakes · Easily confused · Key vocabulary · Related
You’ll learn to
- Define a web search engine and explain the crawling, indexing, and serving/ranking stages in order.
- Explain how crawlers discover pages and why being crawled does not guarantee being indexed or ranked.
- Apply specific keywords, quotation marks, and operators (site:, minus, filetype:) to refine a search.
- Distinguish labeled advertisements from organic results and evaluate a source's credibility.
- Explain, at a neutral level, how SEO and personalization shape the results you see.
Common mistakes
Believing the search engine scans the live internet when you hit enter.
It searches its own index, a stored, pre-built copy of the web. Crawling and indexing happen in advance; your query is answered from that database.
Treating the top result as the most trustworthy or the 'correct' answer.
Ranking estimates relevance and quality, not truth. The top organic slot is the algorithm's best guess, and the very top of the page may be a paid ad. You still evaluate the source.
Not noticing the difference between ads and organic results.
Ads are labeled (for example 'Sponsored' or 'Ad') and are placed because someone paid. The FTC requires them to be distinguishable; look for the label before you click.
Assuming everyone sees the same results in the same order.
Results are personalized by factors like location, language, and device, so orderings differ from person to person. There is no single neutral list.
Typing long conversational questions and never using operators.
Specific keywords plus a few operators, quotes for an exact phrase, minus to exclude a word, site: to pin a domain, filetype: for a format, usually beat a vague sentence.
Easily confused
Organic result vs. Sponsored result (ad)
An organic result is placed by the engine's ranking of relevance and quality; a sponsored result is placed because an advertiser paid, and must be labeled so you can tell them apart.
Crawling vs. Indexing
Crawling is discovering and downloading pages; indexing is analyzing those pages and storing them in the database. A page can be crawled yet never indexed.
Ranking high vs. Being trustworthy
A high rank reflects the algorithm's relevance-and-quality estimate and can be influenced by SEO; it is not a guarantee that the source is accurate, which you must judge yourself.
Key vocabulary
- Search engine
- Software that finds web pages matching a query by searching its own pre-built index of the web and returning ranked results.
- Crawler (spider/bot)
- An automated program that discovers and downloads web pages, mainly by following links and reading sitemaps, so the engine can process them.
- Index
- The large database in which a search engine stores information about the pages it has analyzed, so they can be looked up quickly for a query.
- Ranking
- The programmatic ordering of indexed pages for a specific query, based on estimated relevance and quality across many factors.
- Organic (natural) result
- A result placed by the engine's own ranking rather than by payment; the opposite of a paid ad.
- Sponsored result / ad
- A result shown because someone paid for placement, required to be labeled so users can tell it from organic results.
- Search operator
- A special term or symbol added to a query to refine it, such as quotes for an exact phrase, a minus sign to exclude a word, site:, or filetype:.
- Search engine optimization (SEO)
- The practice of adjusting a site to help search engines understand its content and users find it; it improves the chance of ranking but guarantees nothing.
- Personalization
- Tailoring result ordering to context such as the user's location, language, and device, so the same query can return different results for different people.
Sources & references
- In-Depth Guide to How Google Search Works — Google Search Central (Google for Developers)
- Refine web searches (search operators and refinements) — Google Search Help
- FTC Consumer Protection Staff Updates Agency's Guidance to Search Engine Industry on the Need to Distinguish Between Advertisements and Search Results — U.S. Federal Trade Commission
- SEO Starter Guide: The Basics — Google Search Central (Google for Developers)
- CRAAP Test (Research Starter) — EBSCO Research Starters
EliExplains lessons are original prose written from the open, credible references above. See Copyright & Licensing.
Researched 2026-08-19
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.

