Cultural heritage
Keeping machine guesses out of the record of truth
A museum wanted to clear a cataloguing backlog with machine-generated metadata. The technology worked; the risk was that a wrong guess written into the catalogue of record is effectively permanent. The architecture followed from that.
A decision driven entirely by how hard it would be to undo.
Problem
A large portion of the digitised collection has images but only minimal descriptive metadata, so it cannot be found. Curatorial cataloguing capacity is a fraction of what clearing the backlog would need. The digital team proposed machine-generated descriptions and subject terms.
Context
A collection management system of record, an ongoing digitisation programme, and a public catalogue used by researchers who cite it.
Current architecture
Collection management system, a IIIF image server, and a public search interface reading directly from the catalogue.
Constraints
- Scholarly: the catalogue is cited in published research; an error propagates into the literature.
- Vocabulary: subject terms must come from controlled vocabularies, not free text, or the catalogue stops being searchable in the way researchers expect.
- Provenance: the source of every assertion must be recorded, which is already the discipline for human cataloguing.
- Capacity: curatorial review time is the scarcest resource in the institution.
Evidence
Each statement placed on the ladder before it was used.
The Getty Art & Architecture Thesaurus is a controlled vocabulary for art, architecture and material culture, organised into facets including Materials, Objects, Activities and Associated Concepts. It holds generic terms only: iconographic subjects and proper names are excluded and live in the Getty Iconography Authority and the other Getty vocabularies.
IIIF provides standardised APIs for describing and delivering images and presentation metadata, and is widely adopted for sharing cultural heritage material.
Machine-suggested terms would be accepted by curators at a useful rate. Untested before the pilot; the pilot existed to test it.
Suggestion plus review is faster than cataloguing from scratch, even after the review time is counted.
Anything written into the catalogue of record acquires the authority of that record, whatever a confidence score says.
Questions that changed the answer
- If a machine-suggested term is wrong and cited in a paper, how is that corrected, and who finds out?
- Can a suggestion be distinguished from a curatorial assertion a decade from now?
- What acceptance rate makes review worth a curator’s time rather than cataloguing directly?
- Does a suggestion layer change what the public catalogue should show?
Options
Including the one nobody wanted to discuss.
Write enrichment into the catalogue
Generate terms and descriptions and write them into the collection management system with a confidence flag.
What it costs: Fast and clean-looking, and effectively irreversible once records are harvested, cited and mirrored elsewhere.
A separate suggestion layer
Machine output lives in its own store, linked to the object, never in the catalogue. Curators promote a suggestion to a catalogue assertion by an explicit act, recorded with their name.
What it costs: Two stores to maintain and a promotion workflow to build. Real complexity, accepted knowingly.
Do nothing and hire cataloguers
Clear the backlog with people.
What it costs: Correct and unaffordable at the scale of the backlog.
Economics
Four horizons, not one estimate.
- Build
- Suggestion pipeline plus the promotion workflow. The workflow was the larger half and the part that makes the rest safe.
- Run
- Inference per object is small; curatorial review is the real run cost and the reason acceptance rate decides the whole business case.
- Change
- Vocabularies and models both change. A separate layer can be regenerated wholesale; catalogue records cannot.
- Exit
- The point of the design. Deleting the suggestion store leaves the catalogue exactly as it was, with no cleanup and no archaeology.
Decision
Generate suggestions into a separate, clearly-labelled layer. A suggestion becomes catalogue data only by an explicit curatorial act, recorded with the curator’s identity and the model version.
Why
The deciding dimension was reversibility, not accuracy. Catalogue records are harvested, cited and mirrored; an error there cannot be recalled. A suggestion layer can be regenerated or deleted entirely, so the institution can be wrong cheaply.
Why not the alternatives
- Write enrichment into the catalogue: It would place unreviewed machine output into a record that carries scholarly authority and propagates beyond the institution’s control.
- Do nothing and hire cataloguers: The right answer at small scale and unaffordable at this one.
Trade-offs accepted
- Two stores and a promotion workflow instead of one system. Complexity accepted, because it is what buys reversibility.
- The backlog clears more slowly than an automatic write would suggest, and honestly so.
Reversibility
The chosen design is cheaply reversible: a suggestion layer can be regenerated or deleted whole, leaving the catalogue untouched. The rejected option was not. Writing into the catalogue of record carries a high reversibility cost, because records are harvested, cited and mirrored beyond the institution. Converting the second into the first is what the architecture buys.
Complexity budget
The second store earned its place. It is the mechanism that makes a wrong machine guess cheap. A vector index and a bespoke review application did not, once the existing collection management system could show suggestions beside the record.
Non-functional requirements
- Every suggestion carries model version, date and confidence
- No catalogue write without a named curator
- Terms resolve to identifiers in the right vocabulary (AAT for object types and materials, the Iconography Authority for subjects), never free text
- The public catalogue never displays an unpromoted suggestion
Decision gate
GO on the suggestion layer, with a pilot on one collection to establish the acceptance rate before any wider rollout, because that rate, rather than model accuracy, determines whether this is worth doing.
Implementation
A suggestion store keyed to object identifiers, terms resolved against the controlled vocabulary, and a promotion action inside the tool curators already use. Nothing new for them to learn.
What the decision was expected to achieve
A measured acceptance rate and curator time per record, which together decide whether the programme scales. A low acceptance rate would be a legitimate reason to stop.
No outcome is claimed. This is an illustrative example, so there is nothing measured to report, and a real engagement would state what happened and how it was verified.
Lessons
- Ask how hard a mistake is to undo before asking how accurate the system is.
- Keep machine output out of the record of truth until a person puts it there, by name.
- Extra complexity is sometimes the right answer when what it buys is the ability to be wrong cheaply.

