The Librarian's First Ledger: On the Card Catalog That Indexed the Web
Before the algorithmic hum of data centers, before the first automated crawler ever stretched its legs, the web was a small and quiet library. Its shelves were sparse, its visitors few. And like any good library, it needed a catalog. This task fell not to a machine, but to a person. Her name was Jean Polly, and her tool was a simple text file named, with characteristic librarian pragmatism, ‘whatsnew.txt’.
In 1992, Polly was the ‘librarian’ for a nascent internet service at the University of Buffalo. The World Wide Web, as conceived by Tim Berners-Lee, was a brilliant system of links, but it lacked a central directory. How would anyone find anything? Polly’s solution was profoundly human-scale. She maintained a manually curated list of new and notable websites, a weekly digest of the entire growing web. To have your site included was to be discovered. It was the ultimate editorial link, the first crawl budget, allocated not by a robot’s logic but by a curator’s discerning eye.
The Manual Crawl
Polly’s process was the antithesis of modern crawling. There was no spidering, no parsing of sitemaps, no discovery through links. Instead, she relied on word of mouth, on Usenet announcements, and on direct emails from the few developers building these new digital ‘spaces’. She would visit each site personally, assess its content and value, and then carefully type its name, URL, and a short description into her ledger. This was a crawl driven entirely by human judgment and effort.
Every entry in ‘whatsnew.txt’ was a conscious decision, a vote of confidence. It was a guarantee of quality and a signal of existence. In modern terms, every listed URL had a perfect ‘crawl priority’ and an infinite ‘crawl budget’ because a human had deemed it worthy. There was no noise, only signal. The entire ‘index’ was a sitemap for the early web, and it was maintained by hand.
This era was brief. The web’s explosive growth made manual curation instantly, hopelessly obsolete. Automated crawlers were not just efficient; they were necessary. Yet, Polly’s ‘whatsnew.txt’ embodies a foundational truth we’ve since automated into obscurity: discovery is, at its heart, an act of human judgment. Our crawlers and algorithms are now the librarians, making billions of decisions per second, but they are built to emulate that original, simple goal—to separate the noteworthy from the noise and guide a user to something valuable. We still build for that first, discerning librarian, even if she now exists as lines of code, tirelessly reading the cards we put in her catalog.
Notes & further reading
A few pages I came back to while writing this: