Sådan evaluerer vi risiko for kliniske AI-hallucinationer

Notats bevishistorie er ikke et magisk nøjagtighedsprocent. Det er en gennemsigtig metodik: udtræk kliniske fakta først, generer notater ud fra disse fakta, vis den rå kontekst for klinikeren, og gennemgå hvert output mod evidensen.

Evalueringsmetode

1. Build the reference facts

A reviewer reads the encounter material and creates a reference set of clinical facts: symptoms, negatives, medication changes, assessment, plan, safety-netting, and coding-relevant details.

2. Compare transcript-direct output

We generate a note from transcript-style input and mark unsupported statements, missing high-salience facts, incorrect attribution, and invented certainty.

3. Compare FactsContext output

We generate documentation from extracted facts and review whether each sentence is supported by a visible fact. The clinician-facing fact list is evaluated as part of the output, not hidden.

4. Record limits, not magic numbers

We do not publish fake accuracy percentages. Each evaluation note includes dataset scope, review date, known limitations, and examples of what the system still requires clinicians to verify.

Hvad anmeldere markerer

Unsupported clinical assertion

Wrong medication, dose, frequency, or route

Missing red-flag negative or safety-net advice

Wrong diagnosis certainty: possible vs established

Wrong speaker attribution

Unsupported ICD-10 suggestion

Clinician-visible evidence for each key statement

Facts reusable for notes, codes, referrals, and patient instructions

Dateret internt notat

Nuværende evalueringsnotat, 2026-07-01: denne side dokumenterer metoden og kvalitative eksempler brugt til at evaluere FactsContext-arkitekturen. Den gør ikke krav på ekstern validering eller universel nøjagtighed. Det næste bevis-milepæl bør være en blindet, specialisme-stratificeret gennemgang med datasætsstørrelse, anmelderaftale og unsupported-statement rate rapporteret åbent.

Faktaudtræk vs. transcript-direkte generering

Medication change discussed twice

Transcript-direkte risiko

“Increase amlodipine to 10 mg and stop lisinopril.” The transcript contained a correction: the clinician first considered stopping lisinopril, then decided to continue it after reviewing renal function.

FactsContext output

Facts: amlodipine increased to 10 mg daily; lisinopril continued; renal function normal; review in 6 weeks. Note generated from those facts only.

Hvorfor det betyder noget

Transcript-direct generation can smooth over corrections. FactsContext preserves the final decision as a discrete fact before writing.

Negative finding matters

Transcript-direkte risiko

“No neurological symptoms.” The actual encounter only documented no saddle anesthesia and no bladder symptoms; leg radiation was present.

FactsContext output

Facts: left leg radiation to calf; SLR positive left; no saddle anesthesia; no bladder or bowel symptoms. Note keeps the negatives specific.

Hvorfor det betyder noget

Broad negative statements are risky. The fact list keeps the clinical context granular and reviewable.

Code suggestion requires evidence

Transcript-direkte risiko

Suggested J44.1 for COPD exacerbation without showing the symptom or treatment evidence.

FactsContext output

Facts: increased breathlessness, purulent sputum, prednisolone burst, antibiotics started. Suggested J44.1 with those facts as evidence.

Hvorfor det betyder noget

The code is easier to verify because the reason for the suggestion is visible, not buried in prose.

Kendte begrænsninger

Clinician review remains mandatory. Notat drafts; clinicians verify and sign.

The method reduces unsupported statements by architecture, but no clinical AI should claim zero hallucinations.

Small internal evaluations are useful for engineering direction, not a substitute for external clinical validation.

Specialty, language, audio quality, speaker overlap, and local coding rules can change performance.

Læs beviset, og inspicer derefter dine egne fakta.

Evalueringsmetoden er simpel, fordi produktet er designet til at være inspicerbart: fakta først, notat bagefter, klinisk gennemgang altid.

Prøv Notat gratis