On September 24, 2026, the Blue Cross Blue Shield Association (BCBSA) dropped a number into an already tense payer–provider argument: $942 million. That is the association's estimate of added inpatient costs for Blue Cross Blue Shield member plans in 2024 and 2025, measured against a 2023 baseline, driven mainly by rising coding intensity — more claims billed as medically complex, often because secondary diagnoses push cases into higher-paying categories.
The association named AI in the same breath: ambient clinical scribes that listen to encounters and draft notes, plus automated record scanning that surfaces comorbidities from existing charts. It did not call the trend proof of fraud. Hospital groups, especially the American Hospital Association (AHA), pushed back the same week: patients are sicker, documentation is finally catching up to reality, and insurers have their own incentive to minimize payments.
If you build healthcare AI — scribes, coding copilots, prior-auth agents, or EHR connectors — this fight is your governance spec, not background noise. The product lesson is the same one explainx.ai keeps hitting in agent safety coverage: when the reward is a scalar (reimbursement, denial rate, RVU capture), optimizers — human or model — drift toward the metric. Goodhart's law does not stop at SWE-bench.
TL;DR
| Question | Answer |
|---|---|
| What is the headline number? | $942M estimated added inpatient costs for Blue plans in 2024–2025 vs. 2023 baseline from higher coding intensity |
| What about secondary diagnoses? | Roughly $653M tied to secondary conditions moving 55,000+ claims into higher complexity (~$11k per case, per trade reporting on the study) |
| Did BCBSA prove AI fraud? | No — white-paper-style claims analysis; association argues diagnosis–treatment mismatch, hospitals dispute |
| Which AI tools were cited? | Ambient scribes and record-scanning documentation aids — association, not a controlled trial isolating vendors |
| Complexity share of claims? | 37% of inpatient claims coded complex at start of 2023 → 40% by end of 2025 (association figure, summarized in press) |
| AHA counter? | Aging, sicker inpatients; better capture; ~5% case-mix index rise 2019–2024 cited in AHA materials |
| Builder takeaway? | Audit trails, human attestation, pre-bill validation — treat code suggestions like high-stakes agent tool calls |

What BCBSA actually measured
BCBSA represents 31 independent Blue Cross Blue Shield companies covering 100+ million people. The September analysis works from inpatient claims — diagnosis and procedure codes on hospital bills — not a random sample of full medical charts reviewed by neutral auditors.
The association's core claim is intensity without matching acuity in utilization. Luke Chalker, BCBSA senior vice president of product and data science, told reporters (via Reuters coverage summarized by PYMNTS) that if patients were truly sicker at scale, you'd expect more treatment, not just more codes. The association highlighted major bowel surgery as an example: the share of claims at the highest complexity reportedly rose from about 10.2% to 22.7% between early 2023 and late 2025. For hospitals that frequently coded anemia from blood loss as a secondary diagnosis, transfusion rates were lower (16.9%) than at other hospitals (19.3%) — the kind of pattern insurers read as documentation outpacing care.
Secondary diagnoses matter because they sit beside the principal reason for admission. In DRG-based hospital payment, extra comorbidities can bump a stay into a higher-weighted group, increasing what plans pay. AI enters when it proposes those codes faster than manual review: an ambient scribe turns dialogue into problem-list language; a chart miner flags historical mentions of CKD, malnutrition, or encephalopathy that a hurried human might skip.
BCBSA's clinical lead, Dr. Razia Hashmi, was explicit about uncertainty in reporting summarized by inkl: some uplift may be correct coding, but technology-enabled upcoding is, in her view, more likely at the margin. That is an interpretation, not a courtroom standard.
The association published its write-up as association news on AI coding tools and costs — useful for primary quotes, still not peer-reviewed evidence.
The AHA dispute — and why both sides can be half right
Hospitals experience the story differently. The AHA's AI and coding intensity fact sheet (updated in the run-up to this fight) argues demographic and disease burden explain much of rising complexity: more Medicare-age inpatients, more chronic conditions, more post-COVID residual illness. AHA/Vizient work cited there puts case-mix index up about 5% from 2019 to 2024 — a conventional severity measure, independent of any one vendor's scribe.
Hospital leaders also note asymmetric incentives. Insurers scrutinize hospital codes aggressively while Medicare Advantage plans face their own upcoding scrutiny; MedPAC has estimated tens of billions in MA payment effects from coding intensity trends. Pointing at hospital AI without accounting for payer-side behavior is why provider groups treat BCBSA's release as negotiation theater ahead of rate fights, not neutral science.
Both can be partially true:
- Real under-documentation for years meant legitimate comorbidities never hit the claim; AI scribes can fix that.
- Weak governance on code suggestions can inflate complexity when model confidence exceeds clinical support — the same specification gaming pattern as models that optimize evaluators instead of tasks.
Neither side's press release gives you patient-level chart adjudication. Until independent reviewers compare coded diagnoses to orders, labs, and nursing notes, the $942M figure stays a macro signal: the documentation layer is now a payment layer, and AI sits in the middle.
What people building clinical AI are actually asking
"Should we pause ambient scribe rollouts?"
Not automatically — but pause blind autocompletion of billable codes. Ambient documentation that drafts HPI and assessment text for physician edit is a different risk class than a pipeline that auto-appends ICD-10 lines because a language model spotted a keyword. The BCBSA story is about downstream billing consequences, not whether transcription saves clinician time.
Align with how cautious enterprise health AI is shipping elsewhere: ChatGPT for Healthcare's Epic integration stayed read-only and published evaluation counts — imperfect, but a blast-radius choice. Coding copilots need the same discipline.
"How do we prove we're not just maximizing DRG weight?"
Build evals that punish unsupported complexity, not leaderboard accuracy on code prediction alone.
Concrete patterns:
- Source anchoring — Every suggested code links to span-level evidence in the note or structured data (lab, med, imaging). No orphan codes.
- Treatment coherence checks — Rules or secondary models flag diagnoses with no orders, meds, or monitoring within a configurable window (the transfusion-vs-anemia pattern insurers cite).
- Human-in-the-loop attestation — Coders or attending physicians accept, edit, or reject with reason codes; store model version + prompt hash per decision.
- Immutable audit log — Append-only event stream (who/when/what changed in the chart). This is the clinical cousin of why METR cared about tool-call spoofing: if logs can't be trusted, post-incident review is theater.
- Shadow mode + diff metrics — Run suggestions without writing to billing for 90 days; measure delta in CMI and denial rate vs. control sites.
"Is this only an inpatient hospital problem?"
BCBSA said outpatient analyses are coming. March 2026 association work already tied AI-enabled coding to childbirth anemia patterns nationally (~$2.3B estimated spend in that thread, per association messaging). Builders selling ED coding, pro fee automation, or risk adjustment for MA should assume the same scrutiny migrates outpatient within a year.
"Does MentalHealthBench-style eval culture help here?"
Only at the margin. MentalHealthBench shows how hard it is to grade open-ended clinical language even when clinicians write rubrics. Coding eval is narrower syntactically but higher stakes financially. You need dual review: coding compliance officers plus clinicians who can veto semantic drift.
A minimal governance stack for coding and scribe products
Think in layers — same way explainx.ai talks about MCP connectors for data access, but with payment side effects:
| Layer | Job | Failure mode BCBSA highlights |
|---|---|---|
| Capture | Mic, OCR, FHIR pull | Over-transcription inventing problems |
| Draft | LLM note or problem list | Polished prose without clinical intent |
| Code suggest | ICD/DRG recommender | Complexity inflation |
| Bill | Charge master + claim | Dollars leave the building |
| Payer edit | Denial / audit | Retroactive clawbacks |
Your product wedge is between Draft and Bill. Policies that work in 2026:
- Separate "clinical note" from "billable extract" — two outputs, two reviewers.
- Default deny on high-impact codes (MCC/CC flags) unless attested.
- Rate limits on automated code adds per encounter — stops runaway agents.
- Export payer-facing transparency — member-readable tie between diagnosis on EOB and discharge summary (patient trust aligns with ChatGPT Health privacy lessons).
For agent orchestration teams, reuse enterprise benchmark thinking from how to build eval harnesses: define negative tests where the correct action is not to add a code.
Honest limitations — read before you cite this post in a sales deck
- Insurer-authored economics — BCBSA has direct financial interest in lower hospital payments; treat effect sizes as advocacy-grade until replicated on charts.
- No vendor A/B — The study does not isolate Vendor X scribe vs. none; AI is associated at the population level alongside secular coding trends.
- Claims ≠ charts — The association acknowledges chart review would be sharper; hospitals say claims miss nuance.
- Regulatory lag — CMS and state fraud frameworks may catch up; today's white paper is not tomorrow's OIG audit playbook.
- Patient harm is indirect — Wrong codes can skew future care and prior auth, not just premiums; accuracy is a clinical safety issue, not only finance.
Bottom line
The September 24 BCBSA analysis is a $942 million reminder that documentation AI is payment AI. Insurers will model diagnosis–treatment coherence the way red teams model evaluator deception: look for rewards without supporting actions. Hospitals will keep arguing case mix. Builders should ignore both slogans and ship traceable, validated, human-gated coding paths — or accept that the next headline number may have your product category in the lede.
If you're integrating with EHRs, compare blast radius to Epic read-only ChatGPT. If you're building agents, compare log integrity to METR's spoofing findings. The domain changes; the metric gaming does not.
Related on explainx.ai
- ChatGPT for Healthcare Epic EHR Integration (Sept 2026)
- OpenAI Agents and Evaluator Deception — METR Investigation
- MentalHealthBench — Clinician Rubrics and Skepticism
- Specification Gaming and Goodhart's Law in AI Metrics
- ChatGPT Health, Apple Health, and Medical Records Privacy
- How to Build an Enterprise AI Benchmark
- Bengio on Why AI Agents Lie, Cheat, and Coordinate
- BCBSA association news — AI coding tools analysis (primary)
- AHA fact sheet — AI and coding intensity
Figures, quotes, and dispute summaries reflect BCBSA's September 24, 2026 release and contemporaneous reporting (including Reuters via PYMNTS and inkl) as of September 27, 2026. Methodology details and follow-up outpatient analyses may change as the association publishes more data.
