Alan Turing Institute · Data Study Group

The CHIRPdb dataset

A structured database of maritime incident reports from two investigating bodies - the UK Marine Accident Investigation Branch (MAIB) and the US National Transportation Safety Board (NTSB) - hosted for you as a standard PostgreSQL database. This page orients you: what you have, how it was built, how the tables fit together, and - most importantly - how far to trust each field. The data dictionary is the per-column reference.

What you have

Read access to a running PostgreSQL database, the source report PDFs, and the paperwork to make sense of both. No access to CHIRP's application, pipeline or APIs; everything you need is in the four pieces below.

The database

A standard PostgreSQL instance your IT service hosts. Connect with any Postgres client and query it read-only.

This data dictionary

Every table and column defined, with the provenance and trust of each field. The companion page.

The schema diagram

The entity-relationship diagram - every table and every link between them, below.

The source PDFs

Every MAIB and NTSB report the database was built from, supplied as zip archives through your IT service. Match a document row to its file with documents.filename.

Working with the database

It is ordinary PostgreSQL - three things worth knowing before you query:

What the dataset is

Each published report (a PDF, from MAIB or NTSB) is broken into sentence-level text plus the report's own stated incident particulars. MAIB reports are additionally linked to structured incident data - occurrences, vessels, affected persons - drawn from the MAIB Data Portal spreadsheets; NTSB has no equivalent spreadsheet source, so that layer is MAIB-only. An analysis pipeline then extracts findings from the text: safety issues, recommendations, and named entities with every mention resolved. Every finding is traceable back to the exact sentences of the source report. CHIRP's SHIELD taxonomy - 113 causal-factor codes in 24 categories under four layers - is loaded, and every passage is embedded so codes can be suggested, but no code has yet been matched to a finding.

documents.jurisdiction separates the two corpora: 'UK' for the 796 MAIB reports, 'US' for the 456 NTSB reports - 1,252 documents in all. The pipeline has been run over all 1,252, so the database is populated end to end; what is absent is listed under What is not included, and the one substantive gap is that missing SHIELD link.

What is actually populated

The database is populated end to end - every pipeline stage has run over the whole corpus. Five tables are empty, for three different reasons; know which before you build on them.

Populated
  • Reports & text - documents, authors, sentences, and the report-derived document_particulars / document_vessels. Both corpora, split by documents.jurisdiction.
  • Spreadsheet incident data (MAIB only) - occurrences, vessels, affected_persons, affected_persons_ppe, the report↔occurrence link on documents.occurrence_id, and the vocabularies taxonomy_terms / ppe_items. No NTSB document carries an occurrence_id.
  • Passages & findings - passages / passage_sentences (one passage per finding), safety_issues (11,254), recommendations (4,127), recommendation_organisations and the organisations they resolve to. Written by the record pass over all 1,252 documents.
  • Narrative entities - narrative_entities (545,649) and narrative_entity_mentions (1,267,349), every mention resolved to an entity and located by character span.
  • Embeddings - passage_embeddings (one vector per passage), shield_code_embeddings (452), passage_embedding_queue drained. Behind the match_shield_codes suggestion function.
  • Seeded vocabularies - full SHIELD taxonomy (shield_codes / shield_code_categories / shield_layers), narrative_entity_types (14), embedding_models, embedding_strategies.
Empty
  • passage_shield_codes - the missing link. Taxonomy loaded and every passage embedded, but no code matched to a passage yet.
  • analysis - a spare table for a future per-passage pass. Nothing writes it.
  • legislation, safety_issue_legislation - the legislation-attachment feature is not implemented.
  • author_identifiers - nothing populates it at this stage.

How the data was produced

The data flows through pipeline stages, and the tables mirror them. Provenance is preserved as a chain of links, not copied data: any finding traces to its passage, the passage to its sentences, each sentence to its document and position.

Stage 1
Ingest

A report PDF - MAIB or NTSB - is registered as a documents row and each page is rasterised to an image; the report's own stated facts go to document_particulars and document_vessels, and documents.jurisdiction records which corpus it belongs to.

Stage 2
Extraction

A vision LLM reads each page image and returns its text and structure directly; deterministic passes then run on top (normalise metadata, rejoin broken lines, repair lossy pages), stored as ordered sentences. Lifecycle: queued → extracting → extracted → verified.

Stage 3
Spreadsheet enrichment

MAIB Data Portal data is loaded independently: occurrences → vessels → affected_persons, classified via taxonomy_terms, linked to MAIB reports through documents.occurrence_id. (MAIB only - NTSB publishes no comparable dataset.)

Stage 4
Record pass

Each report is read for the records it publishes - safety issues (findings and general safety lessons, split by record_type) and recommendations - each written with its own passages row. A recommendation's addressee is resolved into organisations. Run over all 1,252 documents.

Stage 5
Entity resolution

A pass reads each document's sentences for the named things it talks about, writing one narrative_entities row per distinct thing and one narrative_entity_mentions row per mention, each located by character span and typed by how it refers.

Stage 6
Embedding

Every passage is encoded under BAAI/bge-large-en-v1.5 (1024 dimensions) into passage_embeddings, and each SHIELD code under four strategies into shield_code_embeddings. These feed the match_shield_codes function - but nothing has yet accepted a suggestion.

How the tables fit together

Two incident-data sources are kept apart and linked, not merged: the report PDF's own account lives in document_particulars / document_vessels (both corpora); the MAIB Data Portal spreadsheet's account in occurrences / vessels / affected_persons (MAIB only); the two join through a nullable documents.occurrence_id FK - several reports can point at one occurrence (e.g. an interim then a final). Tree-shaped classifications (event type, ship type, place on board, injury, deviation) are normalised into taxonomy_terms. The full entity-relationship diagram:

100% Scroll to pan, or hold Ctrl/Cmd and scroll to zoom.
CHIRPdb entity-relationship diagram showing all tables and their relationships

The data dictionary is authoritative for table and column definitions; the diagram is the map, the dictionary the legend.

Where each value comes from - and how far to trust it

This is the point of the dictionary. The database shows you structure; it cannot tell you which values are source fact, which a model produced, and which a human has since confirmed. Every column is labelled with one of the five below - the first two both mean "the source said so", and differ only in which source: a report PDF, or the MAIB spreadsheet.

Source Source-document field - from the report PDF, MAIB or NTSB (content or file metadata).
MAIB MAIB Data Portal field - from the MAIB spreadsheet extract, on the tables it feeds (occurrences, vessels, affected_persons and their vocabularies). UK incidents only.
Pipeline Pipeline-derived, deterministic - computed on ingest, no model involved.
AI-U AI-generated, unverified - produced by an LLM or NLP model, not confirmed by a human.
AI-V AI-generated, human-verified - model output an analyst has confirmed.
System System metadata - surrogate keys, timestamps, status flags, generated search columns, seeded reference data.

A slashed label - Pipeline / AI-U - marks a column whose extractor is deterministic with a model fallback.

Two rules run throughout. Row-level verification: tables with an is_verified flag hold AI-generated content whose status varies per row - treat is_verified = true as AI-V, everything else AI-U. Sentence text is extracted by a vision LLM reading each page image, with deterministic passes on top, so it is AI-U, qualified by the parent documents.status (verified = analyst-confirmed).

Practical upshot: the analysis layer is fully populated, so AI-U content is widespread - all sentence text, every safety issue and recommendation, every narrative entity and mention, and the embedding vectors. None of it has been analyst-verified: no row carries is_verified = true and no document is at status = 'verified'. Treat every AI-U column as machine output no human has checked.

Known limitations

Read these before drawing conclusions from the data.

What is not included

Five tables are empty, for three different reasons:

Also absent: pipeline observability data (LLM traces, extraction logs) and any application authentication data.

Row counts

As of 28 August 2026. Seeded vocabularies are loaded reference data.

TableRows
documents1,252 (796 MAIB, 456 NTSB)
document_particulars1,252
document_vessels1,252
authors22
author_identifiers0 (empty)
occurrences9,063 (MAIB only)
vessels9,830 (MAIB only)
affected_persons3,040 (MAIB only)
affected_persons_ppe924 (MAIB only)
taxonomy_terms410 (seeded)
ppe_items9 (seeded)
sentences341,057 (293,454 MAIB, 47,603 NTSB)
passages15,381 (13,892 MAIB, 1,489 NTSB)
passage_sentences15,939 (14,382 MAIB, 1,557 NTSB)
analysis0 (empty)
passage_shield_codes0 (empty - no code matched yet)
safety_issues11,254 (10,038 MAIB, 1,216 NTSB)
safety_issue_legislation0 (empty)
legislation0 (empty)
recommendations4,127 (3,854 MAIB, 273 NTSB)
recommendation_organisations3,658 (3,342 MAIB, 316 NTSB)
organisations802 (710 companies, 92 classes)
organisation_identifiers86
narrative_entities545,649 (446,735 MAIB, 98,914 NTSB)
narrative_entity_mentions1,267,349 (1,032,182 MAIB, 235,167 NTSB)
shield_codes113 (seeded)
shield_code_categories24 (seeded)
shield_layers4 (seeded)
narrative_entity_types14 (seeded)
embedding_models1 (seeded)
embedding_strategies4 (seeded)
passage_embeddings15,381 (one per passage)
shield_code_embeddings452 (113 codes x 4 strategies)
passage_embedding_queue0 (drained)
Full data dictionary →
Every table and column, with per-field provenance and constraints. The reference to keep open while you query.