In September 2026, the Blue Cross Blue Shield Association (BCBSA), a group representing US health insurers, published a white paper showing that the share of hospital inpatient claims classified in higher-severity billing categories rose from about 37% in early 2023 to about 40% by the end of 2025. Using 2023 as a baseline, the paper estimated this shift produced roughly $942 million in additional spending over two years.

For major bowel procedures, the paper's central example, hospitals where billed severity rose the most did not show higher ICU utilization or transfusion rates than other hospitals. BCBSA suggests this mismatch may be linked to AI tools that extract billing diagnosis codes from clinical documentation.

The American Hospital Association (AHA), however, counters that multiple factors are at play—an aging patient population, rising chronic disease, and a shift of milder cases to outpatient care. Tracing the figures in the white paper, along with the source of the adoption rate BCBSA cites as evidence of AI's spread, reveals a boundary between what the claims data actually shows and what is being inferred from it.

AD

Inpatient claims classified as "more severe" rose from 37% to 40%

In US inpatient care, hospitals are widely paid a fixed amount per Diagnosis-Related Group (DRG). For the same procedure, if a secondary diagnosis is recorded that qualifies as a complication or comorbidity (CC), or a more severe major complication or comorbidity (MCC), the claim is classified into a higher-paying DRG and reimbursement increases.

Even if a patient's actual condition hasn't changed, adding a previously unrecorded diagnosis can shift the billing classification. What the white paper analyzes under the term "coding intensity" is exactly this kind of rise in billed severity.

The paper is formally titled "Hospital Coding Intensity Analysis: Major Bowel Procedures." Written by Chris Birkmeyer, David Wennberg, Keith Kamons, and Luke Chalker, it analyzes claims data from Q1 2023 through Q4 2025. BCBSA states that its member Blue Cross Blue Shield plans cover roughly one in three Americans.

For the paper's central case—major bowel procedures (DRG 329–331)—claims with the most severe MCC designation rose from 20.2% to 22.7%, while claims with neither a CC nor MCC fell from 36.6% to 32.8%. The paper estimates this classification shift alone generated $60.8 million in additional cost.

The paper also identifies which secondary diagnoses increased the most. The steepest gains were in diagnoses that are relatively easy to infer from lab values, such as unspecified acidosis (E87.20), hyponatremia (E87.1), and acute posthemorrhagic anemia (D62).

The paper points to possible contributing factors, including ambient listening tools that generate clinical documentation from patient encounters, and systems that surface candidate diagnoses from lab data. Hospitals have been adopting documentation-support AI tools such as Abridge and Microsoft's Nuance DAX Copilot, as well as AI systems aimed specifically at supporting medical coding, such as CodaMetrix, SmarterDx, and Fathom.

It's worth noting, though, that this white paper is not a peer-reviewed academic study. It was produced by an association of insurers—the party that pays medical claims—using its own claims data.

Of the $942 million, only $653 million comes with a documented breakdown

According to the paper's calculations, relative to the 2023 baseline, hospitals classified an additional 55,158 inpatient stays into higher-severity categories. This is said to have produced an average of $11,800 in additional payment per case, totaling $653 million.

The white paper and its accompanying press release describe this as roughly 70% of the total additional cost.

Of the $942 million BCBSA estimated, only $653 million—about 69%—is tied to a specific, itemized breakdown showing how secondary diagnoses pushed claims into higher billing categories. The remaining roughly $289 million has no detailed breakdown in the paper.

Dividing $653 million by $942 million comes to about 69.3%, leaving a gap of roughly $289 million. Both figures come from the same white paper, covering the same time period and the same set of claims.

For that remaining roughly $289 million, the paper provides no specific evidence one way or the other regarding a link to AI. It can be neither confirmed nor ruled out.

According to Becker's Hospital Review, Chalker said at a September press briefing that the $942 million figure represents the portion that could be specifically linked to a mismatch between billed severity and actual treatment provided—not the full amount of spending attributable to the rise in severity classification overall.

In other words, $942 million is an estimate of part of a larger spending increase, and only about 70% of that portion comes with a documented breakdown.

This is not the first time BCBSA has released this kind of estimate.

In March 2026, working jointly with Blue Health Intelligence, BCBSA reported that nationwide spending potentially linked to AI-assisted coding totaled roughly $2.3 billion—about $663 million in inpatient costs and at least $1.67 billion in outpatient costs.

However, the body of that March study specifically documents only $22 million tied to obstetric cases and an increase of up to about 20% in per-member inpatient cost. The methodology behind the nationwide $2.3 billion figure is not explained, nor is it clear whether this represents an annual figure.

The March figure of $663 million and the September figure of $653 million are close, but they come from different analyses using different methodologies covering different populations. Similar numbers alone don't mean they're measuring the same thing.

AD

Why the white paper stops short of blaming AI directly

One of the stronger pieces of evidence in the white paper linking the trend to AI is a comparison of two groups of hospitals for major bowel procedures in 2025.

Comparing the top 25% of hospitals with the largest increases in billed severity against all other hospitals yields the following:

Metric (2025, Major Bowel Procedures) Top 25% by Rise in Severity Classification Other Hospitals
Share Classified into Higher Category 75.6% 65.0%
ICU Utilization Rate 11.5% 13.2%
Transfusion Rate 3.6% 3.9%
Reoperation Rate 1.7% 1.5%
Length of Stay (median) 4.0 days 4.0 days
Anemia Diagnosis Rate 13.7% 9.9%
Transfusion Rate Among Anemia-Diagnosed Patients 16.9% 19.3%

At hospitals in the top 25%, the share of claims classified into a higher category was more than 10 percentage points higher. Yet ICU utilization and transfusion rates were actually lower, and length of stay was identical. The only treatment metric where the top-25% group scored higher was reoperation rate, by just 0.2 percentage points.

Anemia diagnosis rates were higher in that group, while the share of anemia-diagnosed patients who actually received a transfusion was lower.

BCBSA Senior Vice President Chalker said in the press release that if patients were truly becoming sicker, treatment should be increasing too, and that this mismatch may indicate that "what AI is detecting are diagnoses that can be captured for billing purposes—not patients who are actually more severely ill."

Even so, the white paper uses careful language when describing the connection to AI, such as "may be linked," "coincides," and "fingerprints."

The reason is that the paper's body contains no data on whether individual hospitals had actually adopted AI tools.

The case for an AI link rests mainly on four observations: the timing of when billed severity began rising, the concentration of the change among a subset of hospitals, the rise of diagnoses that are easy to infer from lab values, and the lack of a corresponding rise in treatment despite the rise in diagnoses.

But none of this amounts to a direct calculation of how much of the dollar figure AI actually caused.

The table comparison also has limits. It compares different groups of hospitals as of 2025—it is not a longitudinal analysis tracking how treatment and billing changed at the same hospitals before and after AI adoption.

The March study did include results from a check against actual medical records for a portion of cases.

Using commercial inpatient claims from select plans covering about 62 million members, analyzed from April 2022 through March 2025, the researchers found that among hospitals showing a surge in acute posthemorrhagic anemia codes for obstetric admissions, the diagnosis rate rose from 4.0% to 12.3%. Meanwhile, the transfusion rate rose only from 0.8% to 1.2%.

A medical record audit conducted by one plan found that fewer than 20% of cases met clinical criteria for the diagnosis. However, this is the only case cited where medical records were actually checked as verification.

In a fact sheet released July 31, 2026, ahead of the September white paper, the AHA countered that the claim AI has driven up coding intensity is "not supported by evidence."

As background, the AHA cites multiple factors: an aging population, rising chronic disease, and a shift of relatively mild cases from inpatient to outpatient treatment.

According to AHA's analysis, 19% of the increase in hospital costs from 2019 to 2024 was due to treating sicker, more complex patients. A separate analysis by AHA and Vizient found that the case mix index, a measure of patient severity, rose about 5% over the same period.

However, AHA's figures cover 2019–2024, while BCBSA's white paper focuses mainly on additional spending in 2024–2025. The two periods overlap only in 2024, and neither source presents data under matching conditions that would directly refute the other's claims.

Tracing the source of the "60%+ of hospital systems" claim

BCBSA's September press release states that more than 60% of hospital systems have begun using AI coding tools, using this as one premise for the AI connection.

This figure is hyperlinked, and following it leads to a blog post from the AI vendor bam.ai.

That blog post cites a survey page from Oliver Wyman, stating that 63% of healthcare organizations have adopted AI-driven automation in RCM (revenue cycle management—the set of operations spanning patient intake, claims billing, and payment collection).

However, BCBSA's phrase—"more than 60% of hospital systems have begun using AI coding tools"—is not the same claim as the figure cited at the link.

Digging further into Oliver Wyman's public page reveals that the 63% figure is not from an Oliver Wyman survey at all; it originates from a survey by HFMA (the Healthcare Financial Management Association), which the Oliver Wyman page is citing.

Meanwhile, in a separate survey Oliver Wyman itself conducted—of more than 200 decision-makers and 90 practitioners at US healthcare provider organizations, including outpatient facilities—the share of organizations that have adopted AI organization-wide across various RCM functions ranges from 20% to 40%.

In other words, these three figures differ in survey population, in the scope of AI-using operations, and in the definition of "adoption."

BCBSA's phrasing refers to the share of "hospital systems" that have "begun using" "AI coding tools."

The 63% figure cited by bam.ai refers to the share of "healthcare organizations"—including outpatient facilities—that have "adopted" AI-driven automation across "RCM operations overall," not limited to coding, and the original source is HFMA.

Oliver Wyman's own figure of 20–40% refers to the share of healthcare organizations that have "adopted AI organization-wide" across various RCM functions.

These figures alone don't let us determine which adoption rate most closely reflects reality. But at minimum, none of them can reasonably be read as meaning "more than 60% of US hospitals are using AI to assign diagnosis codes."

The March study cited yet another set of surveys.

According to a 2024 survey by ONC (the Office of the National Coordinator for Health IT, part of the US Department of Health and Human Services), about 33% of US hospitals were using generative AI. A separate survey published in JAMIA (the Journal of the American Medical Informatics Association) in fall 2024 found that among 43 nonprofit hospitals, 60% were piloting or had adopted ambient listening, and 45% were piloting or had adopted autonomous coding.

Both figures are cited in BCBSA materials, but their survey populations, scale, and the specific AI use cases they cover differ from the "more than 60%" phrase used in the September press release.

The March study also cited vendor-supplied adoption case studies, including a claim that McLaren generated $11.3 million annually after adopting SmarterDx, and that Mercyhealth saw a 5.1% revenue increase after adopting Arintra.

AD

Insurers have their own coding and AI-review controversies

The question of how diagnosis coding affects payment amounts has long been contested on the insurer side as well, not just among hospitals.

In a January 2025 report, MedPAC (the Medicare Payment Advisory Commission, an advisory body to Congress) provisionally estimated that differences in coding intensity in Medicare Advantage (MA)—the federal health program for seniors run through private insurers—would add $40 billion to 2025 plan payments.

Under MA, government payments to insurers are adjusted based on the diagnoses recorded for enrollees, so an increase in diagnosis codes can increase insurer revenue.

MedPAC's estimate of the payment gap does not attribute it to a single cause, but AHA cites this figure while describing it as "overpayment from upcoding."

On March 11, 2026, the US Department of Justice announced that Aetna agreed to pay $117.7 million to settle allegations under the False Claims Act that it submitted inaccurate diagnosis codes for MA enrollees.

The settlement covered allegations involving chart reviews for payment year 2015 and diagnosis codes related to morbid obesity for payment years 2018 through 2023.

Meanwhile, AI tools used on the insurer side to review claims are also being contested in court.

In a class-action lawsuit alleging UnitedHealth used the AI tool "nH Predict" in decisions about post-discharge care coverage for MA enrollees, a federal district court in Minnesota ordered broad discovery on March 9, 2026. Optum has countered that nH Predict is not a tool used to determine coverage decisions.

CMS (the Centers for Medicare & Medicaid Services) also launched a pilot program called "WISeR" in January 2026 across six states, using AI-assisted prior authorization for certain traditional Medicare services.

AHA has also criticized insurers for "downcoding"—automatically reducing claim amounts without adequately reviewing medical records.

Between March and September 2026, a string of related developments unfolded: an insurer group's AI coding analysis, a hospital group's rebuttal, and legal proceedings involving insurers' own diagnosis coding and AI-based claims review.

Date Event
March 2026 BCBSA and Blue Health Intelligence publish a coding intensity study focused on obstetric anemia
March 9, 2026 Federal district court in Minnesota orders broad discovery in UnitedHealth's nH Predict lawsuit
March 11, 2026 US DOJ announces $117.7 million settlement over Aetna's MA diagnosis coding issues
July 31, 2026 AHA releases fact sheet on AI and coding intensity
September 15, 2026 HCA Healthcare's CFO tells an investor conference that hospitals are lagging behind insurers in adopting AI for claims processing
September 2026 BCBSA's white paper on major bowel procedures estimates $942 million; press release issued September 24

The remarks by HCA Healthcare CFO Mike Marks were reported by Becker's Hospital Review. Marks also said that administrative costs tied to claims and review are "enormous" for both hospitals and insurers.

According to the same publication, executives on the insurer side have raised similar concerns. UnitedHealthcare CEO Tim Noel said in June that AI-driven revenue cycle management tools are one "factor" in rising healthcare costs, while Aetna Chief Medical Officer Ben Kornitzer described the effect as "broadly inflationary."

Hospitals are adopting AI on the side that generates claims, while insurers are adopting AI on the side that reviews those claims. BCBSA's white paper can be understood as an attempt to analyze the hospital-side shift using claims data held by insurers.

In Japan's DPC system, only 255 of 3,248 codes change classification based on secondary diagnosis

In Japan's DPC/PDPS system, used to calculate inpatient medical costs at acute-care hospitals, secondary diagnoses are likewise one factor that can shift a case into a different diagnosis group.

According to an explainer from the healthcare management outlet Leap Journal, of the 3,248 classification codes in fiscal 2024, only 255 codes had classifications that varied depending on secondary diagnosis. That's down sharply from 1,381 in fiscal 2022.

In other words, the number of points where adding a secondary diagnosis changes the classification fell to less than a fifth over two years.

Even if AI-assisted documentation reduces missed diagnoses, how much that affects billed amounts depends on how much of the mechanism—where adding a secondary diagnosis changes the payment category—still remains.

The controversy playing out in the US points to a need, as AI-assisted documentation and coding support spread in Japan too, to check in advance which added diagnoses actually affect payment amounts.

In the US, this issue is also being discussed in connection with household healthcare costs.

According to KFF's 2025 Employer Health Benefits Survey, the average annual premium for employer-sponsored family health coverage was $26,993, up 6% from the previous year. Of that, workers themselves paid an average of $6,850.

However, neither BCBSA's white paper nor AHA's fact sheet indicates how much of this premium increase is attributable to changes in diagnosis coding.

To turn BCBSA's estimate into stronger evidence for judging the relationship between AI and healthcare costs, at least three additional pieces of information would be needed.

First, what actually caused the roughly $289 million of the $942 million that has no documented breakdown. Second, medical record audit results covering a broader scope than the single insurance plan examined in March. Third, an analysis linking the timing and use cases of AI adoption at individual hospitals to actual claims data.

If changes over the same period could be compared between hospitals that adopted AI and those that didn't, it would become possible to more clearly separate diagnoses that increased because documentation became more thorough from changes reflecting patients who actually became sicker.

Only once that data exists will insurers and hospitals be able to debate, on shared terms, which portion of the AI-driven increase in diagnoses represents legitimate billing—and which portion crosses into overbilling.