The Museum Curator's Hidden Back Room: On the Art of Selective Indexing

Every great museum has a public face: the sunlit galleries where masterpieces are displayed for all to see. But the true heart of the institution, its soul and its memory, lies in the labyrinth of back rooms, storage vaults, and climate-controlled archives. Here, far from public view, resides the vast majority of the collection. A curator’s most critical skill is not merely acquiring art, but deciding what to show, what to store, and what to deaccession entirely. This is a profound lesson for anyone managing a website’s crawl budget.

A crawler, like a museum visitor, has finite time and attention. It cannot possibly process every single object in the collection. The curator understands that to display everything is to display nothing; the sheer volume would overwhelm and confuse, rendering the experience meaningless. They make deliberate, strategic choices. The iconic Van Gogh goes on the wall. The preparatory sketch for that painting is carefully catalogued and stored for scholarly research. The damaged, duplicate, or irrelevant piece is respectfully let go.

We must apply this curatorial mindset to our sites. Our ‘gallery’ is the search engine index, and it should be reserved for our masterpieces: the canonical, valuable, and unique pages that serve a clear purpose for a human audience. Our ‘back room’ is for the supporting cast: the filtered product views, the archived blog posts from 2012, the massive PDF manuals. These are not ‘bad’ pages, just as a storage room isn't full of ‘bad’ art. They are simply not for the main exhibition. They should be accessible, but they don’t need to be highlighted on the main tour.

And then there are the pages we must ‘deaccession.’ These are the broken, duplicate, or thin pages that serve no user or archival purpose. They clutter the collection and drain the crawler’s limited resources. A curator would not keep a crumbling, unsalvageable canvas; it detracts from the integrity of the whole. Similarly, we must be ruthless in removing or excluding such content from the crawl path, using `noindex` or disallow directives as our tool for respectful removal.

This approach transforms crawl budget from a technical constraint into an act of editorial vision. It asks not ‘How can I get everything indexed?’ but ‘What is the most meaningful collection I can present?’ It’s about crafting a coherent, valuable experience for both the automated crawler and the eventual human visitor. By thinking like a curator, we stop seeing our sites as warehouses of content and start building them as thoughtful exhibitions, where every indexed page earns its place on the wall.

Notes & further reading

A few pages I came back to while writing this: