Back to Blog

September 18, 2026

Why Most AI Scribes Fail on Accuracy — and How Cognivolt Doesn't

CONTENT
Every AI scribe can write a clean clinical note. That stopped being difficult two years ago.

What separates them now is whether the clean note is true.

This matters more than it sounds, because AI notes don't fail the way human notes fail. A rushed human note looks rushed, and the reader knows to treat it carefully. An AI note that has dropped a medication is still properly structured, correctly formatted and professionally worded. Nothing on the page signals a problem. It reads as finished, because as a piece of writing it is finished. It's only incomplete as a record.

That's why "just review the output" doesn't work in practice. Review catches errors that look like errors. To catch these, you'd have to hold the whole consultation in your head and compare it against the page, for every note, at the end of every clinic. Nobody does that, and it isn't a discipline problem.

Most AI scribes have simply decided not to solve this. The model writes the note, the note is the product, and the working assumption is that a good enough model produces a good enough note.

Cognivolt is built the other way round. A language model is an excellent writer and an unreliable witness. So it writes, and then what it wrote is checked against what was actually said, before the note is saved.

Here's what that catches.

Reversed denials

Ask clinicians to name the worst thing a note can get wrong and nobody says "a misspelled drug name". They say this: the patient denied thoughts of self-harm, and the note records that he reported them. Or the reverse, which is worse.

It's easier to get wrong than it looks. Take a sentence like hallucinations, delusions, thoughts of self-harm — he denies all of these. The denial lands after everything it cancels. A model reading forward meets three serious findings before it meets the word that takes them back.

Real consultations make it harder. The question belongs to the clinician and the answer belongs to the patient, often several exchanges later. And every good psychiatric note ends with safety-netting — come back immediately if these thoughts appear — a line that names the finding precisely in order to record its absence.

Most scribes don't check for any of this, because checking means comparing the finished note against the recording instead of trusting the model that produced both.

Cognivolt compares them. The findings where a reversal is dangerous — self-harm, risk to others, psychosis, mania, substance use, allergies — are checked against what was said before the note is saved. The check knows a denial can come before or after what it negates, that the answer may sit in the next speaker's turn, and that safety-netting isn't a diagnosis. Where note and recording disagree, the recording wins.

Plans that stop halfway

Read your last AI note from the bottom up.

The Subjective is usually excellent. The Objective is fine. The Assessment reads well. The Plan has the advice and the therapy discussion, and then stops. The drug is gone. The review interval is gone.

There's a structural reason. The plan is the last thing said and the last thing written, so whatever thins out across a long generation thins out there. Everything above it is polished, so nothing marks the place where the writing gave up.

The plan is also the only section of a clinical note that is natively a list: drug, dose, route, frequency, investigations, referrals, safety advice, review date. A list can end early and still read like a finished sentence. Prose can't. That asymmetry is the trap.

Cognivolt doesn't treat the plan as a paragraph to tidy up. It treats it as a set of orders that each have to be accounted for, and it checks the end of the consultation specifically — the part every system is weakest on precisely because it comes last.

Medications

Audit enough AI notes and the errors aren't evenly spread. They gather almost entirely around medication, which is also the part of the note carrying real clinical and legal weight.

The failures we see, repeatedly:

  • The drug disappears. Dictated clearly, absent from the plan.
  • The frequency disappears. Fluoxetine 10 mg instead of fluoxetine 10 mg daily, which reads as once, twice or three times a day to whoever picks up the chart next. That's not a shorter version of the order. It's a different order.
  • The dose changes, and nothing in the sentence looks wrong.
  • The unit changes. Micrograms become milligrams: a thousandfold difference, one word apart.
  • A titration loses half of itself. 5 mg for a week, then 10 mg arrives as 5 mg.
  • A dose attaches to the wrong drug when two medications share a sentence.
  • A range appears that nobody prescribed, reading as a titration that was never ordered.
  • Medication the patient already takes is recorded as started today.

All of those produce notes that read perfectly.

Cognivolt verifies every medication in the recording against the note: the name, and the dose, frequency and duration stated with it. A medication is treated as one indivisible fact, so carrying part of it across counts as an error rather than a summary.

None of this depends on recognising the drug, which matters more than it might seem. There are thousands of medicines, new ones arrive constantly, and every brand name, combination product and national generic is another name a list would need to contain — which is why drug-name lists eventually fail somebody. A prescription has a shape: something named, a quantity, an instruction to take it. That shape holds whatever the drug is called, so a medication Cognivolt has never seen is guarded as carefully as one it has.

Content invented from silence

Point a speech model at audio with nothing in it — a pause, a bumped microphone, an empty room — and it won't tell you the audio was empty. It writes something.

We've seen a silent stretch mid-consultation come back as a complete examination finding, in correct clinical language, with plausible detail, describing something that never happened.

This is the one worth worrying about most, because it doesn't lose information. It manufactures it, and leaves no seam where you might have caught it.

Cognivolt assesses audio for actual speech before transcribing it. Recordings without speech never reach the model, and a transcript that couldn't physically fit inside the speech it came from doesn't reach the note. Fabrication is caught by the arithmetic of the recording, not by hoping the model shows restraint.

Multilingual consultations

Most scribes are built for one language, spoken cleanly, by one person at a time. European clinics aren't that.

A Berlin practice sees Turkish-speaking families. A Brussels consultation moves between French and Dutch without anyone remarking on it. Clinicians across the continent keep the clinical vocabulary in English while the grammar carrying it stays German, Spanish, Italian or Polish.

Single-language systems don't fail loudly here. They fail politely: they take the gist, drop the specifics, and return something shorter and smoother than what was said. Because the output stays fluent, the loss is invisible.

Cognivolt handles the way clinicians actually speak. Dictate in the language you think in and get the note in the language your record requires. Sentence structures that place a denial after what it negates are ordinary grammar here, not an edge case, and the checks above run in the language the note is written in — so a reversal is caught whichever language the consultation happened in.

Scale scores

PHQ-9 nine and PHQ-9 fifteen are different patients. GAD-7 seven and GAD-7 seventeen are different conversations.

Said quickly, in a sentence carrying several numbers, a digit slips. Unlike a reversed denial there's no clue in the wording. The sentence is perfectly correct and only the number is wrong.

Cognivolt checks the scores in the note against the scores in the recording, across the instruments clinicians actually use, and corrects from the source rather than from the model's recollection of it.

The difference, plainly

Most AI scribesCognivolt
Who checks the noteThe model that wrote it, if anyoneChecked against the recording, independently
Reversed denialsNot checkedChecked, in the note's own language
Medication, dose, frequencyTrusted as writtenVerified against what was said
Content invented in silenceModel free to writeBlocked before transcription
The end of the planWhatever the model producedChecked for what was said last
Scale scoresTrusted as writtenVerified and corrected from the source

None of these ask the model to mark its own work. They run against the recording, on every note.

We're not claiming a model that never errs — nobody has one, and anyone selling you one is selling you the demo. The claim is narrower and more useful than that: these errors are known, catalogued and caught before the note reaches you. On the things a clinical note cannot afford to get wrong, Cognivolt is the most carefully checked note in this industry.

If you want the short version of what Cognivolt checks and how, we set it out here: The Most Accurate AI Medical Scribe: How Cognivolt Verifies Every Note.

Try it on your own consultation

Cognivolt is a clinical scribe, decision support and EMR in one, built for all specialties — from psychiatry and psychology to cardiology, dermatology and general practice — with the note formats clinicians actually use, including SOAP, BIRP and DAP.

Questions we get asked

What makes an AI medical scribe accurate?

Not transcription quality. Every current model transcribes well. What matters is whether anything checks the finished note against what was actually said. Most scribes don't. Cognivolt does, independently of the model that wrote it.

Why are AI scribe errors so hard to spot?

Because the writing is good. A note missing a medication is still correctly structured and professionally worded, so there's no visible defect to catch your eye. Errors that look like errors get found in review. These don't.

What is the most dangerous mistake an AI scribe can make?

A reversed denial — a patient's denial recorded as a finding, or a disclosure recorded as a denial. It changes the risk assessment, it changes what the next clinician does, and it's invisible on the page.

Why do AI notes lose the end of the plan?

The plan is said last and written last, so anything that degrades across a long generation lands there. It's also the most list-like section, and a list can stop early while still reading like a finished sentence.

Can an AI scribe handle multilingual consultations?

Cognivolt can. Clinicians dictate the way they speak, including switching languages mid-sentence, and receive the note in the language the record requires, with verification running on the note itself.

Does Cognivolt work outside psychiatry?

Yes. It covers all specialties, with note structures appropriate to each rather than one generic template stretched over everything.

Is Cognivolt GDPR compliant?

Yes. Our security page covers how data is handled, and our DPA and BAA set out the contractual position.