← Back to Blog

How AI-assisted cataloguing meets archival standards: confidence scores, audit trails and human review

Bound volumes shelved in a wooden bookcase, each shelf carrying its own reference label
fig. 1Bound volumes shelved in a wooden bookcase, each shelf carrying its own reference label

The validator’s question

You’re the specialist archivist who knows exactly what standard a system must meet. You built the schema. You spot the material-type mis-classification that would embarrass the service. You’re the person who raises the objection that matters — often after the meeting, once you’ve seen the detail.

So when someone proposes an AI-assisted cataloguing tool, your question isn’t “how fast is it?” Your question is: can I trust it not to invent facts, not to claim certainty where there’s ambiguity, and not to fail at scale in ways that reflect on me?

It’s the right question. “Confident-wrong AI” — systems that hallucinate and present fiction as fact — is the single most disqualifying risk for archival work. A mis-described date, a smoothed-over gap, a description of what the system expects rather than what’s actually there — that’s not just incorrect, it’s professionally damaging. It undermines the integrity of the catalogue, exposes the service to accreditation risk, and erodes public trust in the collections you’re accountable for.

Why standards are the access story

Archival standards exist to enable access. ISAD(G) structure, EAD3 exports, Dublin Core interoperability, controlled vocabularies, provenance documentation — these aren’t bureaucratic obstacles. They’re the infrastructure that makes collections findable, trustworthy, and usable across systems and institutions.

When an AI-assisted tool meets those standards and shows its working, it unlocks something powerful: the ability to catalogue at item-level depth you never had capacity for, surface collections that have sat unprocessed for decades, and do it in a way that stands up to scrutiny. But only if the trust mechanics are visible, auditable, and honest about limits.

This is how it works.

The trust stack: observable, not a black box

Per-word confidence scoring

Every word of every AI-generated description carries a numerical confidence score. During pilot testing at a county archive service, the system processed a handwritten Victorian will — difficult Copperplate script, legal terminology, 19th-century conventions. Average confidence across the description was 0.93. When a single word scored 0.05, the system flagged it for human review.

A county archivist told us: “I’d rather it flag than make it up.”

That’s the design principle. When the AI cannot read a word with confidence, it signals uncertainty rather than guessing. When it cannot determine a fact from the source material, it leaves the field empty. Better a gap than a fiction.

Confidence scoring isn’t hidden metadata — it’s visible in the review interface, exportable in the audit trail, and part of the record’s provenance. You can filter by confidence threshold, prioritise low-confidence items for closer review, or accept high-confidence drafts with minimal amendment.

Reasoning transparency

The system doesn’t just describe — it shows why. During the same pilot, the AI autonomously extracted entities and relationships from the handwritten will: familial links (“daughter of”), discrepancies between stated birth and baptism dates, the authority that had stamped the document. None of this was explicitly prompted; the system inferred structure from context.

A validation specialist noted: “Very interesting that it’s understood that from the context.” Another observation: “I don’t know if I’d have pulled out the authority that had stamped it, but the AI did.”

The reasoning is logged. You can see which fields were AI-generated, which were human-edited, what model produced the draft, and what confidence it carried. The audit trail is exportable and forms part of the item’s descriptive provenance — a requirement for archives working to TNA guidance.

The human review gate

Nothing is published until you accept it. Machine-generated vs human-reviewed status is always visible. The AI drafts; the archivist decides.

This isn’t a workaround for an imperfect system — it’s the architecture. Condition can only be assessed from physical examination of the object, not a photograph. Material type (manuscript vs typescript, vellum vs paper) requires specialist judgement. Subject classification, sensitivity decisions, and public vs internal field visibility all sit with the qualified archivist.

The tool compresses the cost of first-pass description. You spend your time on verification, context, and judgement — not data entry.

Standards-aligned structure and exports

The system maps to archival standards by design:

  • Hierarchical description: genuinely multilevel fonds/series/sub-series/item structure, with collective ISAD(G) description at every level.
  • EAD3 finding aids: nested multilevel exports with <c> component structure, authority-linked <controlaccess>, and ISAD(G)-mapped descriptive fields.
  • Dublin Core: OAI-PMH 2.0 harvesting endpoint for interoperability with aggregators and institutional repositories.
  • CALM CSV export: shipped, for migrating out of CALM, including former-reference concordance so the legacy reference stays visible alongside the current one.
  • SPECTRUM-mapped object cataloguing: for museums and mixed collections.
  • Configurable controlled vocabularies: resolve against LCSH, FAST, Getty AAT/TGN, VIAF, Wikidata, or your own authority files.

You set the standard; the AI drafts to it. Org-level prompt rules, custom fields, and public/internal field visibility controls mean the tool adapts to your practice, not the other way round.

Automated sensitivity flagging

In one pilot session, the system flagged an image with a content warning — “I think there’s a dead body here” — and automatically marked the record as restricted, before the reviewing archivist had noticed the issue themselves.

GDPR-relevant personal data, potential safeguarding concerns, and culturally sensitive material are surfaced for human review. The AI describes what it sees; you decide access conditions.

What we don’t claim

Honesty is the validator’s currency. So here’s what this system will not claim:

Not autonomous. AI-assisted, human-reviewed. The archivist stays in control at every step.

Not infallible. Confidence scores and audit trails exist because the AI will mis-read, mis-classify, or miss nuance. One pilot participant put it plainly: “Even if the AI isn’t entirely accurate, it’s still going to be better than having nothing.” The tool accelerates first-pass description to a reviewable state — not to a finished, unverified publication.

Not TNA-approved or accreditation-certified. The National Archives does not approve or certify systems. We align outputs to TNA guidance and produce standards-ready exports; you assess fit against your accreditation requirements.

Not “fully ISAD(G)-compliant” in every nuance. Standards alignment is a spectrum and context-dependent. We map to ISAD(G) descriptive structure and produce compliant EAD3 exports; edge-case hierarchy rules and local practice variations remain your responsibility to configure and verify.

Condition and material-type assessments from images are provisional. The AI describes what it sees in a photograph; only physical examination confirms conservation status, and only a qualified archivist confirms material type. These fields are flagged for verification before publication.

EAD3 is an export format, not an OAI-PMH metadata format. The OAI-PMH 2.0 endpoint serves Dublin Core only. EAD3 finding aids are available as downloadable exports, not via the harvesting protocol.

Proof: handwritten Victorian will, zero amendments

The hardest test in the pilot: a handwritten will from the 1870s, Copperplate script, legal phrasing, testator and beneficiaries named, property and bequests itemised, witnesses listed. The system produced a complete item-level catalogue record — title, description, date, creator, scope and content, associated names, reference number — in under a minute.

The reviewing archivists made no amendments.

“The AI read faster than Gemma and I could,” one participant said. The other noted: “Considering the handwriting, it did amazingly well.”

Speed matters when you’re facing a backlog of 5 million items with a team of three. But speed without correctness is worse than doing nothing. The handwritten-will test demonstrated both: throughput and standards-aligned accuracy, validated by specialists who knew what to look for.

Invitation: try to break it

Validators want to test edge cases. We want you to.

Ingest a bundle of your hardest material — damaged documents, multilingual records, complex hierarchies, sensitive content, unusual formats. Run the AI’s first-pass description. Inspect the confidence scores. Check the audit trail. Review the EAD3 export structure. Test whether public/internal field visibility behaves as documented.

If it breaks, we want to know how. If it mis-classifies, we want the example. If the learning curve is confusing, we want that feedback. One pilot participant described their role as “a good canary in the coal mine for things that people will do wrong with the system.” That’s the scrutiny that makes the tool better.

Trial access is available for institutional evaluation. No cost, no obligations, no minimum volume. Upload your test corpus, review the output, and report what you find.

Standards-aligned cataloguing at scale is possible. But only if the trust mechanics are transparent, the limits are honest, and the archivist stays in control.


See example exports — EAD3 finding aids, Dublin Core records, CALM CSV mappings, and audit-trail JSON at help.archivers.ai/exports

See it in action — before/after cataloguing examples showing confidence scoring, human review gates, and audit trail output at archivers.ai/demo

Planning a cataloguing or digitisation project?

Archivers.ai sits in front of your existing repository or CMS, clears digitised backlogs faster, and exports into the systems you already use. Now in early access — join the waitlist and we’ll onboard you personally.

Join the waitlist Book a demo

Updates from Archivers.ai

Take the next step

Process your backlog. Deliver public access.

Archivers.ai is the AI processing and access layer for archives and museums — now in early access, with every account personally onboarded.

EU-resident AI · human-review gated · standards-native exports