◎ OireachtasDB

Evidence

What was measured, and what wasn't

The deliverable is the evidence, not the demo. Everything on this page was computed against the live API and is reproducible from src/audit/ and src/eval/.

Corpus, as built

1,200

sittings, 2020-01-21 to 2026-09-24

320,568

speeches parsed, zero parse failures

99.6%

speeches resolved to a member URI

2,004

speeches detected as Irish-language

What didn't work

A skip-paged query silently loses records. Paging a year window returns exactly debateCount rows — but that set can contain a duplicate URI and omit a real record. On 2024 it returned a Seanad record twice and dropped one Dáil sitting entirely. The count check passes while the data is wrong, and which record vanishes changed between two runs 40 minutes apart.

I nearly published 0% speaker linkage for 1919–2024. The list endpoint inlines a speakers array that looked like a free corpus-wide answer. It is populated only for records reprocessed since ~May 2026 — sections declare speakerCount: 3 and return speakers: []. The XML gives 84–99.8% linkage in every decade.

My Irish-language estimate was inflated twofold because my Irish word list contained an, do and go — all ordinary English words. Corrected, the First Dáil went from 58% to 35.7%.

The party palette failed its own test. The colour first chosen for Renua sat at ΔE 6.5 from Aontú — two indistinguishable teals. It was caught by the checker, not by eye, and replaced by search.

Full log: Exploration Log.md. Full audit: Coverage Audit.md.

Known limitations

  • Written answers are not in this corpus. They are not carried in the debate XML at all — only a note saying they are published separately — so they need a separate /questions ingest.
  • Committee debates are out of scope for v1; the corpus is Dáil and Seanad chamber records only.
  • The corpus runs from 2020. The pipeline is date-parameterised, so widening to the 1998 Good Friday Agreement boundary is a config change.
  • Drift is baselined but not yet measured — that needs a second observation separated by time, so no drift figure is claimed.
  • Irish-language detection is a token heuristic, not a trained classifier, and has not been validated against hand-labelled speeches.

Rhetoric–vote alignment

The measure behind the "Rhetoric & votes" button on each member page. Experimental, and the coverage is the first thing to understand about it.

The question — do members vote the way they speak — is only answerable where a member actually spoke in the debate their vote belongs to, and where that vote can be tied to a specific question. Both cuts are steep, and they are properties of parliament rather than of the data:

StageVotesShare
Votes cast in the corpus 186,260100%
Member also spoke in that debate section 15,3128.2%
… and the vote is attributable to a specific question 7,3483.9%

For roughly 92% of votes the member never spoke in the debate. That is ordinary parliamentary behaviour, not missing data. The second cut exists because 1,493 of 1,896 section-linked divisions carry only the subject "Amendment put:" with no amendment number recorded anywhere on the division, while a single committee-stage section can hold fifteen divisions and discuss forty-two numbered amendments. Aligning them by order was tested and does not hold, so those pairs are stored unscored rather than guessed.

Why not sentiment. Valence and stance are different things. "The Minister has been asleep at the wheel on housing" is strongly negative and says nothing about how that member voted. Position is therefore read by a three-step cascade — an explicit move of the question, which is true by construction; declarative stance phrases ordered so that negations resolve first, because "I cannot support" and "I do not oppose" both defeat naive keyword matching; and zero-shot entailment where neither fires. Tone is scored separately and reported separately.

Why the whip breakdown sits beside the headline. Voting against your own stated view because your party requires it is normal, and a raw alignment rate mostly measures party discipline. Each scorable vote is therefore also labelled by whether the member went with their party's majority, giving four cells rather than one number. Independents — with 17,383 votes cast and no whip — act as a natural control: if this measure tracks anything real, unwhipped members should behave differently from whipped ones. That comparison is a live test the measure can fail.

Not yet established. No alignment figure is quotable until the scorer is scored. A hand-labelled gold set with bootstrap intervals is the gate, and it is not built yet; the mover-anchored slice gives an optimistic bound in the meantime, because members who move amendments state their positions unusually explicitly. Treat every number behind that button as provisional until this paragraph says otherwise.