✳ BEHIND THE CRAWL

The web doesn't
discover itself.

Before a page can appear in a search result, something has to find it. Meet the quiet explorers that turn a tangle of links into a searchable world.

FIELD GUIDE
05 MIN READ
FIRST, A SMALL DISTINCTION

The spider in this playground is a visual metaphor. Real crawlers are software. They request and read resources; they do not highlight, enlarge, tilt, or knock words off a website. Those effects make the act of discovery visible.

01
DISCOVERY & THE URL FRONTIER

Start somewhere.
Then follow a link.

A crawler needs a starting point: a known URL, links from a previous crawl, or a sitemap supplied by a website. It keeps discovered addresses in a queue often called the URL frontier.

That queue is carefully managed. URLs may be normalized and checked against addresses already seen. A scheduler decides what to visit next, balancing freshness, priority, and the load placed on each host.

SEED URL→URL QUEUE→FETCH + PARSE
NEW LINKS RETURN TO THE QUEUE ↵
02
ROBOTS.TXT & POLITENESS

Read the room.

Before fetching pages, a well-behaved crawler checks the site's /robots.txt. This file tells matching crawlers which paths they may request. Responsible crawlers also pace requests, limit concurrency, and back off when servers struggle.

# A simple robots.txt example
User-agent: *
Disallow: /private-drafts/

Sitemap: https://example.com/sitemap.xml

A robots rule controls crawling. It is not a password, and it does not reliably keep a URL out of search. A noindex directive can tell supporting search engines not to index a page, but the crawler must be allowed to fetch it to see the directive. Protect confidential content with authentication.

03
HTTP & RENDERING

Ask. Receive. Inspect.

The crawler makes an HTTP request, identifying itself with a user agent. The server responds with a status code, headers, and usually a body. A successful HTML response gives the crawler the source of a page.

200 OKA resource is available to process.
301 / 308A permanent redirect points to a new address.
404 / 410The resource is missing or gone.
429 / 503Slow down or try again later.

Some crawlers also render JavaScript in a browser engine to see content created after the initial HTML arrives. Others only parse the response. Rendering takes extra resources, so discovery and rendering may happen at different times.

04
PARSING & LINK EXTRACTION

A page is full
of possibilities.

The response is parsed into a structure the crawler can inspect. It can extract headings, visible text, metadata, canonical hints, and links. Relative links are resolved against the page's address to produce usable URLs.

<a href="/journal/curiosity">
  Follow your curiosity
</a>

// Resolved against https://example.com
https://example.com/journal/curiosity

New eligible links are deduplicated and returned to the frontier. Crawling policy determines which hosts, paths, file types, and link relationships are followed. This repeating loop lets a crawler explore a connected part of the web.

05
PROCESSING & INDEXING

Finding is only
the beginning.

Crawling collects resources. A search engine's indexing systems then analyze eligible content, group duplicates, choose representative URLs, and organize signals so pages can be retrieved efficiently.

For example, an inverted index maps a term to documents that contain it. When someone searches, a separate retrieval and ranking process evaluates relevant candidates. A page being crawled does not guarantee that it will be indexed or appear for a particular query.

“curiosity”
/journal/curiosity/notes/asking-better-questions/reading-list
06
FRESHNESS & RECRAWLING

The web never
stands still.

Pages change, links break, and new documents appear. The scheduler revisits known URLs, using signals such as past changes and server responses to decide when another visit is useful.

Conditional requests can avoid downloading unchanged content. A server may respond with 304 Not Modified when a stored version is still current. There is no universal schedule: crawl frequency depends on the crawler, the site, and available resources.

Discover. Fetch. Understand. Repeat.
A little exploration makes a big web usable.

KEEP EXPLORING · PRIMARY SOURCESGoogle Search Central — How Search worksGoogle Search Central — Introduction to robots.txtGoogle Search Central — GooglebotLet the spider loose