The Weaver of Cobwebs: On How a Spider Brought Us the Searchable Web

Before the great indexes of Google and Bing, before the term ‘crawler’ was commonplace, the World Wide Web was a quiet, un-mapped territory. You found treasure not by typing a query, but by following a trail of blue, underlined links. This was the world as Tim Berners-Lee conceived it: a web of associations, a researcher’s dream. But a web without a comprehensive index is like a library with no card catalog, its knowledge locked away in unconnected rooms. The individual might have a map of their own corner, but the grand atlas of the entire web was yet to be drawn.

This is where our historical figure emerges, not from a Silicon Valley garage, but from the halls of MIT. Her name was Martijn Koster, a Dutch software developer whose contribution to web discovery was so fundamental we barely utter his name. In the early 90s, Koster created the web's first well-known, well-behaved crawler. He didn’t call it a crawler, though. He called it the Internet Spider, and with it, he began to systematically weave a web of his own over the nascent internet.

Koster’s work was born of necessity. He was maintaining ALIWEB, an early search engine that relied on site owners manually submitting their own ‘sitemap’-like index files. It was a protocol, a politeness. But it was slow, and human nature being what it is, compliance was patchy. The web was growing too fast for a manual process. So Koster built the Spider to go out and actively discover what was there, to find pages that had never been formally announced. This was the paradigm shift: from waiting for the web to register itself to actively exploring its expanding edges.

But Koster’s true legacy isn’t just that he built a spider; it’s that he taught it manners. He authored the now-famous document, ‘Guidelines for Robot Exclusion,’ which introduced the robots.txt protocol. In an act of profound foresight, he created a way for webmasters to politely tell automated agents like his spider which parts of their site were off-limits. This wasn’t just about preventing overloading servers—it was about establishing a code of conduct for a new, automated frontier. He recognized that the act of crawling wasn’t a right, but a privilege that required respect for the spaces being visited.

When we talk about crawl budget and discovery pathways today, we’re standing on the foundation Koster and his peers laid. The questions we grapple with—how to ensure our important pages are found, how to guide crawlers efficiently, how to manage server resources—are merely modern echoes of the problems he identified at the very beginning. Koster’s spider wasn't just a program; it was a philosophical statement. It declared that for the web to be truly useful, it needed an automated, yet respectful, method of self-discovery. Every time a modern crawler respects a robots.txt file, it’s following a protocol written by the first weaver of the web’s great index.

Notes & further reading

A few pages I came back to while writing this: