AI Reality Check: AI in the Real World
Chapter 6 of 7: Carriers and other businesses are increasingly using AI. What are they learning?
A NextAgency Resource
Last Updated: April 17, 2026
Carriers and other businesses are in a race to deploy AI. How’s that working out for them?
Previous chapters in the AI Reality Check series have explored what opportunities and perils AI poses. We’ve discussed how ignorant AI is of its own weakness. Have these considerations impacted how artificial intelligence is being deployed in the real world? Is it possible to be both smart and fast when deploying AI or is it a choice of one or the other?
The answers needed research, not discussion. So this chapter is a departure from others in the series. Claude and ChatGPT researched AI deployments across a variety of industries. They presented their findings to Alan Katz, NextAgency’s co-founder and their co-author of the AI Reality Check series. Together they selected and arranged the examples into a helpful narrative.
When AI Works, It Works
These deployments are narrow, well-bounded and very impactful. They don’t replace human judgment; they support it.
JPMorgan’s Contract Analyzer: In 2017, JPMorgan launched COIN — software that reviews commercial loan agreements and extracts key clauses from standardized documents. It replaced 360,000 hours of annual legal review time. Not interpreting law. Not making legal calls. Just fast, accurate extraction of structured data. JPMorgan Software Does in Seconds What Took Lawyers 360,000 Hours, Bloomberg, Feb. 28, 2017.
Mastercard Fights Fraud at Scale: In 2024, Mastercard deployed a hybrid AI system combining generative models and graph analytics to detect payment fraud and reconstruct card-theft patterns from dark web leaks. It significantly improved detection rates and reduced false positives — without removing human oversight. Mastercard Uses AI to Combat Card Fraud, PYMNTS, May 28, 2024.
Viz.ai Speeds PE Treatment: At TriHealth, implementing Viz.ai’s solution to automatically flag suspected pulmonary embolism on CT scans and alert care teams was associated with a reported 74% reduction in in-hospital mortality risk in one real-world study. Physicians still made all treatment decisions. Shapiro et al., presented at American Venous Forum 2024; BusinessWire, March 7, 2024.
Microsoft DAX Copilot Assists Clinicians: Microsoft’s DAX Copilot listens to physician-patient conversations and drafts clinical notes for doctor review and sign-off. Adopted across hundreds of health systems by 2024–25, it reduces the time physicians spend on post-appointment documentation, a key source of burnout, while doctors approve every note. A Year of DAX Copilot: Healthcare Innovation That Refocuses on the Clinician-Patient Connection, Microsoft, September 2024.
When AI Fails, It Fails
When companies push AI too far and without the proper controls, AI doesn’t just fail. It can fail spectacularly.
UnitedHealth’s Claims Review Bot: In 2023, a class-action lawsuit alleged UnitedHealth used an AI model called nH Predict to deny rehab claims by estimating how long patients should need care based on historical data. The AI had a reported 90% overturn rate on appeal. UnitedHealth says doctors were involved. Plaintiffs say denials were rubber-stamped. The case is still being litigated. UnitedHealth Sued Over AI That Allegedly Denied Claims, Reuters, Nov. 14, 2023.
Cigna’s Algorithmic Autopilot: Also in 2023, ProPublica exposed that Cigna doctors were approving bulk claim denials through an automated system, sometimes spending seconds per case. A class-action followed, as did California’s Physicians Make Decisions Act, effective January 2025, requiring physician review for any medical-necessity denial. How Cigna Saves Millions by Having Its Doctors Reject Claims Without Reading Them, ProPublica, March 2023.
Air Canada’s AI Creates Policy: In 2024, a Canadian tribunal ordered Air Canada to honor a refund promised by its AI chatbot — even though the policy didn’t exist. The court rejected Air Canada’s argument that the chatbot was a “separate legal entity” responsible for its own statements. If your AI says it, you said it. Air Canada Ordered to Honor AI Chatbot Refund, Global News Canada, Feb. 2024.
McDonald’s AI Goes Rogue: After piloting IBM’s AI voice-ordering system in more than 100 U.S. drive-thrus, McDonald’s ended the trial in June 2024 amid widely publicized misorders — including viral videos of the system repeatedly adding hundreds of Chicken McNuggets to a single order. McDonald’s Ends AI Drive-Thru Test with IBM, Restaurant Business, June 2024.
AI Is What AI Is
Google’s Deepmind Misdiagnoses: In 2023, Google demoed Med-PaLM 2 with claims of expert-level medical accuracy. Independent testing found unsafe responses in roughly 18% of cases. AI failure or marketing exaggeration? You decide. Google’s Med-PaLM Gave Dangerous Medical Advice, STAT News, Sept. 2023.
Google Definitely Overpromised This Time: Also in 2023, Google released a highly produced video showing Gemini reasoning across text, images, and video in real time. Journalists later found it was staged — still images fed one by one with prewritten prompts. Don’t blame the AI; blame the people trying to pull a fast one. Google’s Gemini AI Demo Was Staged, TechCrunch, Dec. 2023.
Volkswagen’s Billion-Euro Misfire: In 2020, Volkswagen launched Cariad to build a unified AI-driven operating system for all 12 VW brands. By 2025, after years of delays and multi-billion-euro losses, it had become one of the auto industry’s most prominent and costly software misfires. Multiple financial and technology outlets, 2024–25.
The “AI” That Wasn’t: In 2025, the DOJ and SEC filed parallel charges against Albert Saniger, founder and former CEO of Nate, Inc., for alleged securities and wire fraud. His shopping app promised users they could “buy what you wish in a tap” — powered by proprietary AI. According to prosecutors, the actual automation rate was essentially zero percent. Contractors in the Philippines and elsewhere were manually completing every order. Saniger allegedly raised $42 million before the scheme collapsed, leaving investors with near-total losses. He faces up to 20 years if convicted. DOJ and SEC charges, April 9, 2025; Fortune, TechCrunch, CBS News.
AI Unchecked, Consequences Follow
Vendors offering unsupervised AI are promising convenience. That’s not always the result.
Fake Law, Real Consequences: In 2023, plaintiff’s lawyers used ChatGPT to draft a federal court filing in the Southern District of New York containing fabricated case citations — complete with invented case names, docket numbers, and fake judicial opinions. Sanctioned $5,000. Widely cited as the first high-profile federal case of its kind. Mata v. Avianca, S.D.N.Y., June 22, 2023.
The Price Goes Up: In March 2026, the 6th U.S. Circuit Court sanctioned two Tennessee attorneys a total of $116,315, jointly and severally, for filing briefs containing more than two dozen fake or misrepresented citations — currently one of the largest AI-related legal sanctions on record. The court called it “an abuse of the adversary system.” Whiting v. City of Athens, 6th Cir., 2026 WL 710568; Dinsmore & Shohl, March 2026.
Books That Don’t Exist: In May 2025, a syndicated summer supplement published in the Chicago Sun-Times and Philadelphia Inquirer recommended books that do not exist — real authors paired with hallucinated titles. The author used AI to help write the section without fact-checking the output. Both papers pulled the section and apologized. AP and multiple outlets, May 2025.
Official Advice, Wrong Answer: NYC’s MyCity chatbot, positioned as a trusted source on city regulations for small businesses, repeatedly gave factually wrong and sometimes illegal guidance — on tips, scheduling, housing, and other topics — without independent legal review of its outputs before publication. The city defended it as a pilot. It stayed online. The Markup, March 2024; CIO.com, December 2025.
The Pattern
I asked ChatGPT and Claude to identify the patterns in these examples, something that AI is generally good at.
1. AI Succeeds When It’s Supervised
Every successful deployment shares three characteristics:
Narrow scope (e.g., extracting loan clauses, flagging fraud patterns)
Clear data inputs (structured documents, defined targets)
Human oversight at key decision points (e.g., underwriters, customer service reps)
The Lesson: AI prepares, humans decide.
2. AI Fails When It Makes Judgments
In failed deployments, AI was used to:
Make autonomous decisions (e.g., denying medical claims or buying houses)
Interpret human needs or risks (e.g., treatment plans, market timing)
Operate on bad or overly simplified training data
The result? Legal, financial, or reputational damage.
3. Hype Consistently Outpaces Reality
Across sectors, AI boosters tend to overpromise and underdeliver:
Overstate capabilities (Google, VW, Nate)
Downplay limitations and error rates
Mislead about “AI” when it’s just automation
4. Lies or Ignorance?
Failures often come down to AI’s inability to verify what it confidently presents as facts:
It can sound confident while being wrong (Air Canada)
It can be trained on bad data and not know the difference (NYC MyCity)
It can mimic decision-making without understanding context (Cigna, UnitedHealth)
What This Means for Life and Health Insurance Agencies
These patterns aren’t abstract—they predict which AI tools will help your agency and which will hurt it. Before adopting any AI system, ask:
Does it make decisions or prepare them?
Can I verify its outputs against source documents?
What happens when it’s wrong?
Who’s liable—the vendor or me?
Bonus question: Is the data handled securely, with privacy maintained?
The Reality Check
There’s a tendency among AI boosters to promise a revolution while delivering results that are incremental. Incremental progress is still progress and over time it can become revolutionary. When listening to those describing the wonders of artificial intelligence, your first question would naturally be, “What does this mean for me?” Your second question should be “What’s in it for them?”
AI is a power tool in a world that thinks it’s a magic wand. It can work wonders when used properly—with safety gear, clear purpose, and your hand on the off switch. If you demand AI perform miracles you’ll be disappointed. If you harness its abilities—and subject them to knowledgeable human review—you can do great things with it. Knowing its strengths and understanding its weaknesses increases the odds you’ll use AI and not fall for the hype. As the AI put it, AI isn’t dangerous because it’s evil. It’s dangerous because people think it’s smarter than it is.
In the next—and final — chapter in this series, the AI and I will offer thoughts on how life and health insurance agencies can responsibly deploy artificial intelligence in their businesses.
Postscript:
This AI Reality Series is based on blog posts from August 2025. In that version of this chapter, we noted Zillow’s experience in 2021 with an AI implementation gone bad, very bad. With all the new examples we found, that story didn’t make the cut for this chapter, but it is too good to lose. So here is a memorable tale of what happens when even smart companies imbibe too much AI-flavored Kool-Aid.
Zillow’s $380 Million Mistake: In 2021, Zillow shut down its home-flipping business after its AI overpaid for thousands of houses. The model couldn’t keep up with market shifts. When prices cooled, Zillow lost at least $380 million in one quarter and 2,000 people were laid off. Zillow Quits iBuying After Huge Losses, CNBC, Nov. 2021.
Learn More:
To find other chapters in the AI Reality Check series as well as information on other topics of importance to life and health insurance agencies, please visit the Resource Center.
Author Information:
All chapters in the AI Reality Check for life and health agencies were co-written by NextAgency Co-Founder Alan Katz, Claude AI and ChatGPT.