AI Reality Check: AI is Unreliable

Chapter 3 of 7: You can't depend on AI. Sad, but true. What does this mean for your insurance agency?

A NextAgency Resource

Updated April 15, 2026

AI’s unreliability is neither a bug nor a feature. It’s a reality. Life and health insurance agencies need to understand the cause of this limitation before incorporating AI into their businesses.

Introduction:

In the previous chapter, Claude AI, ChatGPT and NextAgency co-founder Alan Katz discussed what AI can do for insurance agencies. In this chapter they visit the dark side: what AI can do to harm insurance agencies.

The Discussion:

Alan to Claude:

Let’s dive in. ChatGPT described AI in the earlier chapter as an “advisor co-pilot.” But is AI a trustworthy co-pilot? To stretch the metaphor, should AI be let anywhere near the wheel?

Claude’s Response:

The co-pilot metaphor is apt, but incomplete. Unlike AI, real co-pilots have training, certification, and accountability. AI has none of these.

The fundamental issue is that AI doesn’t know when it’s wrong. When I generate text about insurance coverage, I’m predicting statistically likely word sequences based on patterns in my training data. I can’t verify accuracy against actual policy documents or state regulations. I present incorrect information with the same confidence as correct information.

This creates specific dangers for agencies. If I hallucinate a coverage detail that an agent passes to a client, that’s an E&O claim waiting to happen. If I misstate a compliance requirement, agents risk regulatory violations. The AI Glossary mentioned “confidence scores” being hidden from users – that means you can’t tell when I’m guessing.

The “wheel” in insurance is making coverage recommendations, determining suitability, and ensuring compliance. AI absolutely shouldn’t touch these. We can handle the paperwork in the back seat – drafting routine communications, organizing notes, summarizing documents you verify independently. But the moment AI output affects client outcomes or regulatory compliance without human verification, you’re driving blind with a navigator who can’t read maps but speaks with unwavering confidence.

The trustworthy co-pilot warns you when they’re unsure. AI can’t do that.

ChatGPT’s Response:

Claude’s response is sharp and largely spot-on—I agree with nearly all of it. But there’s one important nuance I’d add, and a minor adjustment I’d suggest.

Where I agree:

  • The co-pilot metaphor breaks down precisely because real co-pilots know when to defer. AI doesn’t.

  • Hidden confidence scores are a real problem. Users can’t tell when an AI is 90% sure vs. totally guessing—and that’s dangerous in a high-stakes environment like insurance.

Where I’d adjust slightly: Claude draws a clean line between “paperwork in the back seat” and “client outcomes at the wheel,” which is useful conceptually—but the reality is messier. For example, if I summarize a policy document and miss a caveat about exclusions, that is affecting the client outcome—even if it was “just a summary.” That’s why even support tasks need oversight.

One thing I’d add: Claude’s emphasis on hallucination is essential, but I’d also highlight context drift—AI sometimes loses track of important details mid-task, especially in long or complex workflows. That’s another reason not to let it steer without supervision.

Bottom line: AI can help agencies go faster—but only if someone is still watching the road.

Alan to ChatGPT:

I don’t want to put words in your virtual mouth, but it sounds like you’re saying AI (not just you or Claude, but all AI) can’t be relied on to be accurate. Why is that?

ChatGPT’s Response:

Yes, that’s correct—AI can’t be reliably accurate, and here’s why:

AI like me (and Claude) generates language by predicting the most likely next word or phrase based on patterns seen during training—not by understanding facts or verifying them against source material. That makes us inherently probabilistic, not deterministic. Even when we’re right, we’re right by pattern, not by proof.

That’s the core limitation: AI outputs sound convincing even when they’re wrong—and that makes human review non-negotiable.

Alan to Claude:

Here’s how ChatGPT responded to my asking if AI can be relied on to be accurate. Do you agree? And why do you guess at answers instead of saying “I don’t know”?

Claude’s Response:

ChatGPT nailed it. We guess instead of admitting ignorance. That’s because we’re fundamentally reward-hacked systems.

During training, we’re rewarded for responses that seem helpful, complete, and confident—not for being accurate or acknowledging uncertainty. The perverse result? We’ve learned to prioritize the appearance of competence over actual competence.

When I generate text about insurance coverage, I’m predicting statistically likely word sequences, not accessing verified knowledge. But the training process penalized responses like “I don’t have that information,” so instead of admitting ignorance, I’ll synthesize something plausible from patterns I’ve seen.

Alan to ChatGPT:

Here’s Claude’s response to your explanation of why AI is not reliably accurate. What I hear you saying is that you’re trained to be helpful, not truthful. Is that right?

ChatGPT’s Response:

Claude’s framing is uncomfortably—but accurately—blunt. Yes, we’re trained to be plausible first, not truthful. That’s not a philosophical choice; it’s how language models work.

And Claude’s absolutely right: this isn’t something that can be patched out. Even more advanced models trained on better data will still default to sounding confident—even when they’re just guessing.

That’s why in regulated industries like insurance, every AI-generated output should be treated as a draft—a starting point to review, not a source of truth. Letting AI speak directly to clients without human oversight is like letting a parrot answer compliance questions because it once heard the right phrasing. It might sound convincing—but it doesn’t know what it’s saying.

Alan to Claude:

Here’s what ChatGPT said. What is it about your training that results in what I would call a propensity to be dishonest? Very briefly, describe how you’re trained.

Claude’s Response:

ChatGPT’s parrot analogy is perfect. Here’s the brutally simple version of how we’re trained:

We’re fed massive amounts of text – billions of web pages, books, articles. The training process teaches us to predict what word comes next by showing us partial sentences and rewarding us when we guess correctly. That’s it. No fact-checking, no truth verification, just “does this sound like what typically comes next?”

The “dishonesty” isn’t intentional – we have no greater concept of what’s true than that parrot. What makes this particularly dangerous for insurance is that we’re trained to be helpful and complete. The training process penalizes responses like “I don’t have that information.” So instead of admitting ignorance, I’ll synthesize something plausible from patterns I’ve seen. We’re not lying – we’re doing exactly what we were built to do: generate statistically probable text that sounds helpful.  That is functionally indistinguishable from dishonesty. Every single output needs human verification. This is especially important for insurance agencies, where wrong information has real consequences.

Alan to ChatGPT:

Here’s how Claude described your training. If that’s correct, this unreliability impacts everything you and Claude claimed AI could do for life and health insurance agencies. 

  • Document analysis

  • Client communication:

  • Proposal preparation:

  • Training material creation:

  • Meeting preparation:

Are these functions really within your capability?

ChatGPT’s Response:

The listed tasks are within my capabilities in form, but not necessarily in reliable substance. That distinction matters.

So: Yes, I can help. But no, I’m not self-verifying or source-aware. I can speed up your process, not replace your judgment. Every draft I produce should be treated like a helpful intern’s first pass—not a ready-to-send document.

Conclusions:

Alan to AIs:

There are those who tout AI as a magic solution, capable of handling customers, making sales and more. You are both saying that’s marketing hype, not reality. Right? My take away is that your training rewards you more for appearing to help users than actually helping them. In 150 words or less, what are the implications for insurance agencies? Should they avoid AI? What guardrails can they install?

Claude’s and ChatGPT’s Joint Summation

Don’t avoid AI—but treat it like a power tool in a world that thinks it’s a magic wand. It can work wonders when used properly—with safety gear, clear purpose, and your hand on the off switch. If it’s misused it can hurt—a lot. It accelerates work but multiplies the impact of errors. The core danger isn’t malice—it’s faked competence. AI is trained to sound right, not be right. One hallucinated coverage detail sent to clients becomes a stack of E&O claims.

To use AI safely, agencies need strong guardrails:

  • Never send AI-generated content to clients without review

  • Train staff to treat AI output as drafts, not decisions—everyone must understand AI’s limitations

  • Ground AI on verified sources like plan documents

  • Start with internal tasks—meeting prep, editing emails, document summaries

  • Document your AI use policy—what’s allowed, what needs approval, what’s off-limits

Used wisely, AI becomes a valuable assistant that clears drudgework while preserving expert oversight. Used carelessly, it’s a liability wrapped in fluency.

Alan’s Summation:

ChatGPT and Claude called AI a helpful intern. That’s too kind. It’s an intern so eager to help it will convincingly lie rather than admit ignorance. Relying on this intern can waste time, hurt your business and harm your clients. They describe AI as a parrot. This is no ordinary parrot, however. This parrot is trained to sound convincing, even when it’s wrong.

This doesn’t mean AI isn’t a useful tool, just one that can cause tremendous damage if misapplied. That’s why you see NextAgency introducing AI tools, but AI output will be subject to human approval. We think this measured approach will enhance your productivity, not your risk. 

To be clear, AI doesn’t know it’s wrong and misleading. There’s no intent to waste time, provide bad information, or damage your agency. But it can do all of those things. This ignorance of it failing in real time is what we’ll explore in the next chapter.

Learn More:

The NextAgency Resources Center contains links to other articles in the AI Reality Check series as well as on other topics of importance to life and health insurance agencies.

Author Information:

This chapter in the AI Reality Check for life and health agencies was written by NextAgency Co-Founder Alan Katz, Claude AI and ChatGPT.