The Janitor and the Geographer: Two Philosophies of Web Discovery

There’s a quiet but firm disagreement simmering beneath the surface of web development, a debate about how we want search engines to find our pages. It’s not a technical specification so much as a philosophical stance, and it can be understood best by picturing two very different figures: the Janitor and the Geographer.

The Janitor operates on a principle of pristine containment. For this approach, the sitemap is everything. It is the master inventory, the single source of truth. The Janitor’s site is a vast, orderly warehouse where every item has a designated spot, and the only way to know what’s inside is to consult the official ledger. Internal linking is kept functional, almost sparse; the primary corridors are clear, but the countless storage rooms remain locked until their number is called from the list. This is a world of explicit declaration. A page exists because it is declared to exist, and it is found because its coordinates have been formally submitted. The crawl budget is spent not on exploration, but on systematic verification of a known universe.

The Geographer, by contrast, distrusts central planning. To them, a sitemap is a useful supplement, perhaps, but it’s a sterile document compared to the rich, organic landscape of the site itself. The Geographer’s work is in crafting a terrain so naturally interconnected that a crawler can’t help but discover its depths. They are obsessed with the flow of contextual links, building a web so dense and intuitive that every path leads to another discovery. A page’s importance isn’t dictated by a list, but emerges from the number of trails that lead to it. The homepage is not just an entry point; it’s a continent, with rivers of navigation flowing into seas of content, each hyperlink a tributary feeding the whole.

The choice between these two figures defines the character of a site. The Janitor’s domain is secure, manageable, and perfectly accountable. If a page isn’t in the sitemap, it effectively doesn’t exist—a useful feature for controlling what gets indexed. But this control comes at a cost. The serendipitous discovery of a deep, rarely-updated but richly linked article is unlikely. The crawler is an auditor, not an adventurer.

The Geographer’s world is alive, dynamic, and often messy. It trusts the emergent intelligence of the network. A crawler becomes an explorer, following scent trails of relevance. The risk, of course, is that some pages—poorly linked or residing in isolated silos—might become forgotten islands, waiting for a geographer’s map that never comes. The crawl budget is an investment in charting unknown territory, with no guarantee of what will be found.

Most of us are neither pure Janitors nor absolute Geographers. Our sites are hybrids, a compromise between the need for control and the beauty of organic discovery. But understanding these two poles is crucial. Are you building a warehouse, where everything is catalogued and requested? Or are you cultivating a landscape, where things are found by following the paths? The answer shapes not just how crawlers see your site, but how humans ultimately experience it.

Notes & further reading

A few pages I came back to while writing this: