A Thousand Cuts: How AltaVista Lost the Web One Click at a Time
Before Google’s PageRank became the dominant logic of the web, there was AltaVista. Launched by Digital Equipment Corporation in 1995, it wasn't just another search engine; it was a monumental achievement in web crawling and indexing. It was the first to truly attempt to devour the entire web, and for a glorious moment, it succeeded. Its crawl was voracious, its index deep and wide. To understand how pages get found, we must first understand how one of the most powerful discovery engines ever built was slowly, methodically, lost.
AltaVista’s initial brilliance was in its raw technical power. Its crawler, Scooter, was a marvel of its time, capable of fetching millions of pages a day. It didn't just skim the surface; it plunged into the depths, indexing the text of documents with an unprecedented thoroughness. For a user, this meant a genuine chance to find a specific, obscure needle in the growing digital haystack. The discovery of a page was a direct function of its existence; if it was online and contained your keywords, AltaVista’s index would likely have it. The crawl budget, as a concept, was simply to crawl everything.
Yet, this technical prowess contained the seeds of its own irrelevance. The very completeness of its index became its user's burden. A search for a common term could return tens of thousands of results, with no intelligent hierarchy to separate the authoritative from the incidental. The engine could find every page, but it couldn’t tell you which one mattered. It was a library with every book ever printed piled in a single, mountainous heap in the center of the room.
The Click That Broke the Crawl
The fatal shift wasn't a technical failure but a philosophical one. As the web commercialized, AltaVista’s owners transformed it from a discovery tool into a “portal.” The clean, singular search box was crowded out by shopping links, news feeds, and entertainment sections. The user’s journey to a result was no longer a direct path but a maze of distractions. Every extraneous click on a portal feature was a click that didn't go to a search result, a silent vote against the purity of discovery.
This portal strategy was a fundamental misunderstanding of the user's intent and, by extension, the crawler's purpose. While Google bet everything on refining the path from query to result, making that journey faster and more accurate, AltaVista buried its magnificent index under layers of noise. It stopped trusting its own core strength—the ability to return a comprehensive, relevant result—and instead tried to guess what else you might want before you'd even found what you were looking for.
AltaVista’s legacy is a cautionary tale for anyone who thinks discovery is purely a technical problem. It’s a human one. The most powerful crawler in the world is worthless if the pathway to its findings is obstructed. It reminds us that how a page gets found is not just about whether it’s in the index, but how the gatekeepers of that index choose to present it. They didn't lose the web because their crawl failed. They lost it because they forgot that every single click is a decision, and they asked their users to make one too many.
Notes & further reading
A few pages I came back to while writing this:
- one area's overview
- The Lighthouse on a Hill of Dirt: What a Forgotten Geocities Page Taught Me About Reach
- a place-by-place guide
- The Patient Art of the Dead End
- a local resource
- The Cartography of Silence: What Library Science Teaches Us About Unseen Pages
- a nearby resource
- a helpful reference
- a regional guide
- a practical rundown
- a useful directory
- a useful directory
- a local resource