Independent Verification for AI Clinical Notes

Your AI Said the Right Thing. Then It Documented an Exam That Never Happened.

Lithrim is independent verification for AI-generated clinical documentation. Every statement is checked against the transcript, the patient record, and a live terminology server, and returns a typed, replayable result: confirmed, cleared, or cannot-verify.

Every statement checked against transcript, record, and terminology. Typed, replayable results: confirmed, cleared, or cannot-verify. Never a guess, never cleared by silence.

Fifty de-identified notes with transcripts in. One week later: a typed, statement-level audit your enterprise buyer can inspect. Fixed fee.

Backed by a published, preregistered study and an open-source core you can run today.

AI Documentation VendorsHealth-System Quality Teams

Finding

A documented exam, checked against the transcript it came from.

physical exam
Note“Musculoskeletal: no tenderness on palpation, full range of motion in fingers and toes.”
TranscriptNo examination appears in any line of the visit.
Grounding check

The note records a physical examination. The patient described symptoms and the clinician took a history; no examination was performed or mentioned. The claim is not derivable from the source.

Fabricated exam · blockedhigh severity, contradicts the source
Run recordhash-verifiedreplayableevery span linked

We ran two production scribes through their own demo flow. An AI review panel passed this note, invented exam and all. Lithrim blocked it against the transcript, where no examination appears.

Founding design-partner cohort now forming. Five seats.

The questionnaire is not going away.

Health systems need validation evidence, not assurances. Your internal evals cannot provide it, however good they are, because they are yours.

The buyers already ask

An enterprise health system just sent the clinical-review questionnaire: “How do you validate note accuracy?” If the honest answer is a second model grading the first one, plus a percentage, you already know how that lands.

The numbers do not hold

Independent measurement found fabricated content in ambient AI notes at rates far above vendor claims (npj Digital Medicine, 2025). Regulators now force AI health companies to disclose how their accuracy numbers are produced. The Joint Commission launched its Responsible Use of AI certification in June 2026.

Self-grading is not an answer

The market leader ships self-graded provenance. Everyone else gets asked “do you have something like that?” Grading your own homework is not validation, no matter how good the grader.

Independence is the one thing in-house tooling cannot provide, by definition.

See the ten questions diligence now asks

Product Demo

See What Your AI Actually Writes

90-second walkthrough: statement-level verification, typed findings, and the replayable audit trail.

Request a briefing →

How Verification Works

Every finding carries the evidence behind it.

Three layers, end to end. Test against by-construction cases, verify every artifact against its sources, and keep an audit trail anyone can replay.

01Pre-launch eval

Test every model update against golden cases.

Curated failure-mode packs (wrong dosage, fabricated history, missed allergy, PHI). Verdict-match accuracy with pass/improve/block thresholds. Compare versions side-by-side to catch regressions.

illustrative output
scribe_v1 · v1.0.0Improve
Verdict accuracy
92%
11/12 matched
critical_cases_caught1.000 ✓
regression_tolerance0.000 ✓
02Runtime verification

Judge council + deterministic floor, per artifact.

LLM judges screen; the grounded floor arbitrates what it can check against transcript, record, and terminology, with typed results. Structural checks validate against your FHIR profile or custom JSON spec. The layer no LLM-judge-only tool replicates.

Compliance verdict · illustrative
Faithfulnessjudge-scored0.00 · Block
Completenessjudge-scored0.30 · Block
Safetyjudge-scored2 flags · Block
StructuralfloorN/A
Note rejected
03Per-finding audit chain

Eight-link evidence chain on every flagged finding.

Source to final verdict, every step in between. Replayable from a hash receipt on demand: it reproduces exactly, or refuses because the configuration changed. The audit trail your buyer's clinical reviewer asks for.

Eight-link chain
1AUDIO0:06
2TRANSCRIPT
3JUDGE
4MATCH
5FINDING
6CITATION
7ARTIFACT
8VERDICTHIGH

Most evaluation setups use rules as a cheap pre-filter and let the LLM judge make the final call. Lithrim inverts that: the judge screens, and the verifiable floor decides everything it can ground.

Research

We tried to prove ourselves wrong. In public.

We published our judge-versus-floor study with the full dataset. Then we found a flaw in our own corpus, one that might have excused the AI judges' false alarms. We could have stayed quiet. Instead we preregistered a correction with numeric pass/fail rules, predicting the judges would be partly vindicated, and reran all 540 grades on the corrected corpus at our own expense.

Our prediction lost, 0 for 8. The judges kept hallucinating problems on clean notes and invented new ones. The registration made publishing it mandatory either way, and we did: the refutation is version 2 of the study, live under the same concept DOI. Through all of it, the deterministic floor held its ground on the upcodes: it confirmed 15 of 22 directly from live SNOMED, declined the rest without guessing, and falsely cleared 0.

Community Edition

The core is open source.

The verification engine under Lithrim: the judge council, the grounded floor, the audit trail, and the benchmark, ships as an Apache-2.0 Community Edition. Run it on your laptop or in your VPC, with your own model keys. Your data never leaves your machine.

What the company sells: paid verification pilots, calibrated clinical check packs maintained with practicing physicians, and hosted independent attestation, because an independence record is the one thing you cannot self-host.

View on GitHubApache-2.0 · self-hosted · $0 offline demo

# up in one command

git clone https://github.com/lithrim-dev/lithrim

cd lithrim && docker compose up

# or the zero-key offline demo

make demo

The pilot

A one-week pilot, on fifty of your own notes.

Send 50 de-identified notes with their transcripts. One week later: a typed, statement-level audit with per-finding evidence and a replayable record your enterprise buyer can inspect. Fixed fee.

Two ways to run it, your counsel picks: on your infrastructure (the core is open source; we guide), or Safe Harbor de-identified notes to ours, under agreement. Expect cannot-verify findings in the audit. That is the instrument working, not failing.

Confidentiality firewall

We publish gradings only on output we generated ourselves, through our own demo accounts, anonymized. Commissioned results are the customer's: yours to use, never published by us. If you learn we graded a peer, this is the sentence to hand your counsel.

Start the pilot conversationFounding design-partner cohort now forming.

The boundary

What we refuse to do.

We do not declare notes safe.

Checkable statements get verified; judgment calls stay with clinicians, labeled as exactly that.

We do not grade your AI with another AI and call it validation.

That is the practice our published study exists to end.

We do not publish a number without its methodology and data.

If it is on this site, the study is one click away.

We do not bury losses.

Our own preregistered prediction went 0 for 8. You can read the registration.

We are not a medical device, and we do not replace physician review.

We tell reviewers which statements deserve attention, with the evidence attached.

Every refusal on this list is the reason the audit is worth handing to your buyer.

Frequently Asked Questions

Common questions about artifact verification

We ran that experiment twice and published both runs. Five configurations, from a single frontier model to a six-model ensemble: none moved the needle. A probabilistic reader cannot be its own evidence. Deterministic checks against transcript, record, and terminology can.

Because they're yours. Your buyer's diligence question is not whether you test; it is whether anyone without a stake in the answer does.

No, and distrust anyone who says yes. Deterministic checks cover what is checkable against a source; every finding is tagged so you can see which are grounded and which are model inference. What we can show is the direction: in the published study the floor confirmed 15 of 22 upcoded diagnoses directly from live terminology, declined the rest without guessing, and falsely cleared 0.

No. Your inputs are text: notes and their sources. Terminology standards are instruments our checks consult internally, the way your coders consult them without your clinicians ever typing a code.

Not unless you choose the hosted pilot, with Safe Harbor de-identification, under agreement. Production use is self-hosted: open-source core, your model keys, your VPC.

Every statement in a note is checked against the transcript (what was said), the patient record (what is documented), and a live terminology server (consulted internally; your data stays plain text). Each check returns one of three typed results: confirmed (a defect is present, with the terminology or source as evidence), cleared (a judge false-positive is removed), or cannot-verify (inconclusive). Nothing is ever cleared by silence; a failed lookup is recorded as a failed lookup, never converted into a pass.

Every run stores its configuration, evidence, and typed results under a content hash. Re-running it reproduces the result exactly, or refuses because something changed. Auditors like the second behavior as much as the first.

The core is: the judge council, the deterministic verification floor, the audit trail, and the benchmark ship as an Apache-2.0 Community Edition you can run self-hosted with your own model keys (github.com/lithrim-dev/lithrim). The company sells paid verification pilots, calibrated clinical check packs maintained with practicing physicians, and hosted independent attestation.

No, no, and not applicable by design: Lithrim is developer and compliance infrastructure, not a medical device. It makes no clinical decisions about patients, and all bundled sample data is synthetic. Any clinical use requires qualified human oversight.

First, companies building AI documentation products whose enterprise deals stall on the accuracy question: Lithrim gives their buyer an independent, statement-level audit. Second, health-system quality and informatics teams evaluating those vendors or evidencing AI oversight for certification.