The Archivist's Unseen Hand: On the Principle of Provenance in Web Discovery
In the quiet, dust-scented halls of a traditional archive, there is a sacred rule that governs every document, every folio, every scrap of paper. It’s called the Principle of Provenance. In its simplest terms, it dictates that the records of a single entity—a person, a government department, a corporation—must be kept together, maintaining their original order and context. To break this chain is to lose the story. The meaning of a single letter is hopelessly obscured if severed from the correspondence that came before and after it.
This isn't just about preservation; it's about intelligibility. Archivists understand that the true value of information is often not in the isolated fact, but in the connective tissue that binds it to its source and its siblings. The structure is the narrative.
Now, step into the digital archive of the web. A crawler is not so different from an archivist on their first day at a massive, disorganized collection. It arrives at a domain and is immediately faced with a universe of pages. Its mandate is to understand, to categorize, to make this information intelligible for the index. How does it decide what matters? How does it discern the important narrative threads from the random clutter?
It looks for provenance. It looks for structure.
When we build websites, we are not just creating individual pages; we are constructing a digital body of records. The internal linking structure we choose is the modern equivalent of an archivist’s meticulous filing system. A shallow, flat site where every page is linked from the homepage tells one story—a story of equal, but perhaps shallow, importance. A deep, hierarchical site, where child pages are linked from their logical parent sections, tells a richer, more nuanced story. It establishes context. It creates a chain of custody for topical relevance.
A crawler following a link from a ‘Research Papers’ section to a specific PDF is following a clear, provenance-rich path. It understands the relationship. It can confidently assign authority and context to that PDF because of its place in the structure. Conversely, an orphaned page, linked from nowhere within its own site but perhaps only from an external blog, is like that single, mysterious letter in an archive with no return address. The crawler may find it, but it cannot truly understand it. Its value and purpose are unclear because its provenance has been broken.
The lesson from the archivist’s world is profound: your site’s structure is not merely a navigation aid for users. It is the primary narrative you present to the crawler, the story that explains how your content is related and why it matters. A clean, logical information hierarchy is the unseen hand that guides discovery, allowing the mechanical archivist to piece together your story, one meaningful link at a time.
Notes & further reading
A few pages I came back to while writing this:
- Oxnard, CA
- The Annotated Path: On the Subtle Signal of Schema's Whisper
- Palmdale, CA
- The Gatekeeper's Dilemma: On the Unseen Weight of robots.txt
- Pasadena, CA
- The Cartographer and the Courier: On the Two Maps of Discovery
- Pomona, CA
- Riverside, CA
- Roseville, CA
- Sacramento, CA
- Salinas, CA
- San Bernardino, CA
- San Diego, CA