The Broken Link Heard Round the House
It was the 404 that started it all. Not a catastrophic, system-wide collapse, but a small, domestic one. My partner, trying to find a recipe for sourdough bagels I’d bookmarked months before, called from the kitchen: “Hey, this link you sent me just goes to a page saying ‘Not Found’.” The complaint was mild, but in my mind, it echoed like a slammed door. It wasn’t just a dead end for her; it was a small failure of my own personal information architecture.
I’d spent years thinking about how massive, faceless algorithms traverse the web, but I’d never applied that same logic to the tiny, shared web of our household. We had our own ecosystem of links: Google Docs for shared grocery lists, a Trello board for home renovation ideas, a sprawling bookmark folder of “Places to Visit.” This was our domain, and I was its primary, and rather negligent, crawler. I had enthusiastically seeded it with pages—the bagel recipe, an article on repotting orchids, a YouTube tutorial on fixing a dripping tap—but I had done nothing to maintain the pathways.
The concept of ‘crawl budget’ suddenly felt intensely personal. A search engine allots a finite amount of attention to a site, deciding which paths are worth following and which are likely dead ends. In our house, my partner’s attention was the crawl budget. Every time she clicked a link I’d shared and found it broken, it was like a bot hitting a soft 404. It wasted her time and, more importantly, eroded her trust in the source. How many broken links would it take before she stopped clicking them altogether? I was squandering the most valuable crawl budget I had.
So, I did what any obsessive would do. I became the dedicated crawler for our home domain. I wrote a small script that periodically checked our shared bookmarks for broken links. It was a crude imitation of the sophisticated bots I wrote about, but its purpose was profound. It wasn’t about indexing for a massive audience; it was about preservation for an audience of two. It was about ensuring that the path to the bagel recipe remained clear, that the tutorial for the leaky faucet would be there when the drip started again.
This small, domestic incident reframed the entire discipline for me. Web crawling, at its heart, isn’t just about discovery for the masses. It’s a form of stewardship. It’s the quiet, ongoing commitment to maintaining the connective tissue of information, whether that information spans the globe or just the distance from the living room to the kitchen. The goal is to ensure that when someone—a stranger across the world or your partner looking for dinner—reaches for a link, the path holds. The bridge doesn’t crumble. The page is there, waiting, faithfully found.
Notes & further reading
A few pages I came back to while writing this: