Home > Health > Expert Contributor

AI Turns 'Good Enough' Into Healthcare's Ticking Time Bomb

By Waleed Mohsen - Verbal
Founder & CEO

STORY INLINE POST

DIA assistant
Waleed Mohsen By Waleed Mohsen | Founder & CEO - Mon, 08/17/2026 - 05:30

share it

"We don't want to know."

I was sure I'd heard wrong, so I asked him to repeat himself.

"We don't want to know," he said again. "That's the truth."

"Oh."

"I'm just being honest."

"No, no. I appreciate it," I said, and sipped my coffee.

This was shortly after I'd founded Verbal. Our product is designed to help healthcare teams audit patient interactions against their clinical protocols, ensuring every interaction meets their organization's standards. As such, I was trying to learn everything I could about how healthcare teams approached quality assurance, clinical adherence and compliance. I'd started meeting with chief compliance officers, chief operating officers, and other leaders at major healthcare organizations, learning how their teams audited interactions, how confident they felt that their rules were actually being followed in the exam room, on the phone and over telehealth.

I wasn't pitching him, just giving context for my questions.

Didn't he want to know how well his organization's rules were being followed? Where the team was doing great and where it could improve? Wouldn't everyone want to know that?

Apparently not.

He must have seen the look on my face, because he took a moment to explain.

It wasn't that he didn't care.

There were just too many interactions, he said. Too many telehealth visits, too many phone calls. There probably were some errors and protocol violations slipping through the cracks, he said, but they simply didn't have the time, staff or resources to check all those interactions, much less fix a bunch of problems.

What was the point of turning on the floodlights and surfacing problems he had no way to fix?

I was surprised by the frankness. But not the sentiment.

He wasn't confessing to some uniquely careless approach. By then I'd learned that, for all their good intentions, few leaders could say with much confidence that their protocols were actually being followed in patient interactions. There was almost always a gap there, a blind spot, an inevitable crossing of the fingers. Everyone just had to make peace with doing what they could and hoping for the best.

"Good enough" had to do. "Good enough" was healthcare's default setting.

But that was a few years ago. Healthcare AI was clearly on the horizon, but still more of a "what if."

Today, AI is more and more central to every aspect of healthcare, from triaging patient calls and drafting clinical notes to managing medication regimens. Soon we'll likely see more and more AI agents actually holding the conversation with the patient directly. And AI is doing those things at a volume no human team could ever match on its own.

So the gap that executive shrugged off a few years ago hasn't gone away.

It's widening.

The Adherence Gap: An Open Secret

The gap that leader described to me isn't rare.

In Verbal's survey of 30 healthcare organizations, spanning hospitals, payers, and telemedicine companies, we found that 70% conduct quality assurance audits on virtual visits once a month or less. Eighteen percent run no formal QA program at all.

Where audits do happen, the sample sizes tell the same story: one regional health plan reviewed just one call per care manager per month, roughly 200 calls against about 46,000 monthly interactions. That covers just 0.4% of interactions. A community hospital did only a little better, auditing 4% of interactions.

And even that thin slice of auditing is limited. The same survey found that 96% of virtual visit QA reviews rely on clinical chart notes alone. They don't even look at the actual interaction via an audio recording or transcript.

Most organizations are only verifying adherence by checking a few notes a month, and those notes are just the clinician's account of what happened.

But chart notes are prone to omission (one study found up to 50% of patient problems discussed verbally in an encounter are never documented in the EHR). They also miss tone and sentiment, so they can easily hide patient confusion or distress. And if anyone downstream even catches an error, it'll only be found weeks after the fact.

For a long time, this gap has been survivable, in the way a small structural crack is survivable in a building that isn't under much load. Errors moved at human speed.

A tired or overwhelmed clinician might deviate from protocol in one or two encounters, and while the odds that encounter fell inside the reviewed 0.4% were slim, the damage stayed contained to those few patients.

An unfortunate situation, to be sure. But inherently limited.

AI doesn't just add another channel to audit, though. It changes the shape of the risk entirely, because it scales both sides of the equation that used to keep the gap survivable: the volume of interactions happening, and the volume of documentation generated about them.

Even if an AI tool is 98% accurate, that 2% gets very scary at scale.

More AI, More Interactions, More Risk

AI agents can now hold conversations with patients, follow up between visits, and in some cases even manage medication autonomously. Every one of those interactions is a chance for something to go wrong. It may even become a pattern, scaled across thousands of interactions.

AI scribes and clinical assistants have documented failure modes of their own: sycophancy, where models agree with illogical or incorrect medical input; hallucination, where they fabricate clinical details or protocol steps that never happened; and a well-documented inability to catch their own mistakes. One study found that un-tuned LLMs could comply with illogical or factually incorrect medical requests up to 100% of the time, and another that non-reasoning language models missed the exact same errors in their own output 64.5% of the time.

But again, think of the scale.

Think of the executive I talked to, who "didn't want to know." His approach was already inadequate for a slower, more human-scale system.

Now we're asking a shrinking, burnt-out workforce to also keep pace with a growing volume of AI-generated interactions and documentation. The volume of things that need checking is scaling exponentially, but the systems designed to check them aren't.

The sort of limited manual auditing we're used to becomes almost meaningless against an AI system generating far more volume, far more quickly.

The gap is widening, the risk is growing, and for many organizations deploying AI tools it may only be a matter of time until it surfaces all at once as a sentinel event.

Bottom Line: 'Good Enough' Won't Do

I still think about that leader who told me, frankly, that he didn't want to know.

At the time, I understood it. Choosing not to look too closely was, in its own way, a rational response to a problem nobody had a real answer for.

But the scale that comes with AI makes this approach untenable.

An error that used to touch one patient can now repeat across thousands of interactions before anyone notices. A blind spot that used to produce the occasional bad outcome can now produce a pattern big enough to trigger a lawsuit, a wave of denied claims, or worse.

Whether it means more rigorous QA, more hiring, or even using AI to QA AI, it's clear that we have to do better than "good enough."

I'm not saying there's an easy fix. But acknowledging the problem is the first step. Prioritizing it is a choice.

The next time I hear "We don't want to know," I won't hear an honest admission of a system's limits. I'll hear a choice.

You May Like

Most popular

Newsletter