The US-based Forecasting Research Institute (FRI) released preliminary verification results on September 22 examining how accurately AI experts and "superforecasters"—people with strong track records in prediction—have anticipated AI's progress.

On math and code-generation benchmarks, actual progress has outpaced not only forecasts gathered in 2022 but also more recent predictions made after ChatGPT's launch.

Meanwhile, many questions about societal effects—such as employment and the economy—have not yet reached their resolution dates.

It can be confirmed that forecasters underestimated the pace of AI capability improvements. But whether that same pattern of error can be directly applied to future predictions about employment and economic impact is a separate matter.

The figures FRI presented need to be viewed by distinguishing between results that have already been finalized, interim progress ahead of a deadline, and future projections.

FRI conducted a cross-analysis of multiple surveys carried out from mid-2022 through August 2026.

The initial participants in the Longitudinal Expert AI Panel (LEAP), which began in 2025, totaled 339 people: 76 computer scientists, 76 AI industry experts, 68 economists, and 119 policy researchers.

Separately, superforecasters selected based on their past forecasting track records also took part.

However, the respondents referred to as "experts" in this analysis are not necessarily the same individuals across all questions.

Also, this release is not a peer-reviewed paper scoring every forecast, but an interim report covering questions the authors judged to be representative.

AD

The AI benchmarks showed the largest gaps

Among the 2022 forecasts, one of the biggest gaps involved the timing of AI reaching gold-medal-level performance at the International Mathematical Olympiad (IMO).

According to FRI, the median forecast among experts was 2030, and among superforecasters, 2035.

FRI determined that AI had actually reached that level by July 2025.

That is five years earlier than the experts' forecast and ten years earlier than the superforecasters'.

Since this 2022 survey was conducted before ChatGPT's public release, it's possible that respondents simply could not foresee the rapid AI development that followed.

However, large discrepancies also appear in more recent forecasts made after ChatGPT's debut.

For the relatively easier Tiers 1–3 of the math benchmark "FrontierMath," LEAP asked what accuracy rate the top-performing AI model would achieve by the end of 2025.

The median forecast was 31% among experts and 30% among superforecasters.

The figure FRI ultimately adopted was 40.7%.

For the code-generation benchmark "LiveCodeBench Pro (Hard)," regarding the highest accuracy rate as of the end of 2026, the median forecast was 14% among experts and 12% among superforecasters.

Yet by May 2026, the top score had already reached 53.8%.

Both figures, however, require caution.

According to FRI, roughly 42% of the FrontierMath problems were later found to contain errors, leading to score corrections.

The 40.7% figure is the uncorrected value FRI used to match the conditions respondents faced at the time they made their forecasts. There remains room for debate over which figure should ultimately serve as the point of comparison.

As for LiveCodeBench, the question asked about the value at the end of 2026, and 53.8% is merely an interim figure as of May.

FRI also notes that some respondents may not have fully understood the evaluation conditions—namely, that models could later be applied retroactively to previously released problems.

While it's clear that AI performance improved faster than forecast, the numbers cannot simply be lined up without considering the evaluation methods and comparison conditions.

Breaking down "missed forecasts" into three categories

The cases FRI examined do not all share the same degree of certainty.

FrontierMath is a question whose deadline has passed and whose result is finalized; LiveCodeBench represents interim progress ahead of its deadline; and the self-driving vehicle example involves an estimated future value.

Question FRI Examined Median Forecast vs. Comparison Value Status as of September 2026
FrontierMath Tiers 1–3, end of 2025 Experts 31%, superforecasters 30%, vs. FRI's adopted value of 40.7% Treated as finalized after the deadline passed. However, this is the uncorrected value from before the problem errors were discovered
LiveCodeBench Pro (Hard), end of 2026 Experts 14%, superforecasters 12%, vs. 53.8% as of May 2026 Actual measured value before the deadline. Final year-end results not yet determined
Share of self-driving vehicles among US ride-hailing services in 2027 Experts 7.3%, superforecasters 2%, vs. FRI's LLM-based estimate of 2.5% Not an actual result but a future projection. The judgment of overestimation is provisional

The classification criteria are straightforward.

Is it a value that has already passed FRI's designated deadline and was adopted as the final result? Is it interim progress observed before the deadline? Or is it a future value estimated by an LLM based on past performance and news coverage?

Because these three categories differ in both subject matter and units, the magnitude of forecast error cannot be simply compared across them.

In particular, the 2.5% figure for self-driving vehicles is not an actual 2027 result but merely an estimate FRI has calculated at this point in time.

Evaluating forecasts midway through their timeline carries another caveat worth noting.

If AI performance surpasses the median forecast early, it can be judged at that point to have been "underestimated."

On the other hand, confirming that progress was slower than expected requires waiting until the forecast deadline arrives.

As a result, interim evaluations that include pre-deadline questions are structurally more likely to surface cases of "underestimation."

FRI itself acknowledges that many questions have not yet produced results and that this approach carries such a bias.

Furthermore, comparing based on median values can obscure the existence of some respondents who predicted the actual outcomes quite accurately.

AD

AI company revenue and broader social adoption are separate matters

Regarding AI company revenue, actual growth may also have exceeded forecasts.

FRI asked about "the highest annual recurring revenue (ARR) recorded by an AI company by the end of 2026."

The median forecast was $20 billion among AI experts, $16 billion among economists, and $25 billion among superforecasters.

Against this, FRI cites reports that Anthropic's annualized revenue as of September stood at roughly $100 billion, pointing to the possibility that this far exceeds the original forecasts.

However, this comparison, too, is not yet settled.

The $100 billion figure is a reported value obtained by annualizing the current revenue pace.

The forecast question, meanwhile, asked about the ARR to be recorded by the end of 2026.

Because this compares the current revenue pace with an accounting figure that will be finalized at year's end, the gap between the numbers cannot be treated as the final forecast error.

Looking at AI adoption across society as a whole reveals a different pattern.

Regarding the share of US working hours assisted by generative AI, the 2025 median forecast was 4% among experts and 3.6% among superforecasters.

The figure FRI adopted as the final result was 5.7%.

This, too, exceeded the forecasts, but not by as wide a margin as with corporate revenue.

Furthermore, the baseline figure of 2% that respondents were originally shown was later revised upward to 3.35% following subsequent research updates.

The fact that respondents made their forecasts starting from a lower baseline than the actual figure may also have affected the outcome.

A surge in revenue at a single AI company is not the same thing as AI usage hours increasing at the same pace across society as a whole.

In real-world tasks, AI's effect was sometimes overestimated

Not all forecasts underestimated AI's progress.

In laboratory tasks related to biology, forecasters actually overestimated AI's effect relative to what was observed.

In the randomized controlled trial FRI cited, the share of science and engineering students who completed all three tasks within the time limit was 5.2% for the group with access to both an LLM and the internet, and 6.6% for the group with internet access only.

The difference between the two groups was not statistically significant.

By contrast, the pre-trial median forecast had predicted a 27% completion rate for the group with LLM access and 12% for the internet-only group.

In other words, respondents anticipated that AI would substantially boost success rates on the experimental tasks, but no such effect was actually confirmed.

That said, the scope of what this experiment measured is limited.

It examined whether specific tasks—three in total—could be completed within a set time limit using the models available at the time; it did not assess AI capabilities across biology as a whole.

Nor does it demonstrate that future, more capable models will be unable to assist with dangerous experiments or similar tasks.

Still, this result does reveal something meaningful.

There can be a substantial gap between achieving high scores on a benchmark and humans actually being able to use AI to successfully carry out real-world tasks.

AI capability evaluations cannot simply be translated directly into real-world productivity or outcomes.

AD

It's still too early to judge the accuracy of long-term forecasts

Among the forecasts FRI collected, many questions concerning societal impact—employment, GDP, serious harms caused by AI, and the like—have not yet produced results.

A separate XPT study FRI published in 2025 analyzed 38 short-term questions from the 2022 forecasts whose results had been confirmed by mid-2025.

That analysis found no statistically significant correlation between the accuracy of short-term forecasts and estimates of long-term existential risk.

This does not mean that long-term dangers posed by AI are small.

Rather, it means there is currently little basis for concluding that experts who missed short-term AI benchmarks are also wrong about long-term societal impacts.

Benchmarks for measuring AI capability are updated rapidly.

By contrast, confirming how much people and businesses actually use AI, how employment and productivity change, and what kinds of benefits or harms emerge each requires different metrics and different timeframes.

Going forward, if questions that reach their forecast deadlines can be scored comprehensively rather than through a few standout cases, it will become clearer in which areas experts and superforecasters underestimated AI's progress and in which areas they overestimated it.

Furthermore, if it becomes possible to verify whether forecast errors on benchmark performance are actually linked to forecast errors regarding real-world work and economic impact, it will become possible to judge more concretely how such forecasts should be used in policy and business planning.