In a September 21, 2026 commentary in The Conversation, Alessandro Di Nuovo and Samuele Vinanzi argued that scenarios in which AI wipes out humanity run into physical constraints, such as the need for equipment and access. They also warned about the danger of humans handing their judgment over to machines. The piece is not a report of new experimental results; it is an essay drawing on existing research and cases. Tracing the mammography reading experiment it cites reveals a gap worth distinguishing: evidence that people can be led astray by wrong advice is one thing, and the concern that judgment erodes over the long term is another.
Physical safeguards are not the same as safety
The authors point to the difference between software capability and the ability to operate real-world facilities. Making a pathogen requires laboratory equipment and the handling of materials, and damaging critical infrastructure requires reaching the relevant systems. Being able to write code does not by itself satisfy those conditions. This is the contributors' own risk assessment, not a research finding that measured the probability of human extinction.
The nuclear sector offers an example of safeguards that limit connectivity itself. In a July 2016 explainer, the US industry group Nuclear Energy Institute (NEI) explained that critical safety and security systems at nuclear plants are isolated from the internet. Using physical network separation or hardware isolation devices, the design keeps office computers from connecting directly to critical control systems.
However, the same NEI document also lists controls on incoming storage media and measures against insiders. Measures that block intrusion from outside networks must be paired with management of the routes into the facility. Nor can protections that an industry group describes for US nuclear plants simply be extended to power grids or airports around the world.
Spelling out physical constraints is useful for assessing danger. But concluding from the mere existence of constraints that every attack path is closed is a leap. What is needed is an assessment tied to conditions, including how much connectivity and control AI is allowed.
What the 27-reader experiment measured
The example the commentary gives of human-side vulnerability is a peer-reviewed 2023 paper by Thomas Dratsch, Xue Chen and colleagues, published in Radiology (DOI: 10.1148/radiol.222176). In the study, 27 radiologists at three German sites read mammograms on a screen described as an AI system.
The test used 50 cases acquired between January 2017 and December 2019. The first 10 cases showed correct suggestions as practice; of the remaining 40, 12 contained incorrect suggestions. In fact, no AI analyzed the images: researchers supplied the classifications and color heat maps highlighting areas of interest. This was an experiment in which people read images, not a predictive simulation run inside a computer.
The outcome measured was accuracy on BI-RADS, a classification of breast imaging findings. As the Radiological Society of North America (RSNA) research summary explains, this classification is not itself a definitive diagnosis. The group averages differed as follows between cases with correct suggestions and cases with incorrect ones.
| Reading experience and number of readers | Cases with correct suggestions | Cases with incorrect suggestions |
|---|---|---|
| Less experienced, 11 readers | 79.7% ± 11.7 | 19.8% ± 14.0 |
| Moderately experienced, 11 readers | 81.3% ± 10.1 | 24.8% ± 11.6 |
| Highly experienced, 5 readers | 82.3% ± 4.2 | 45.5% ± 9.1 |
The figures are group means from the paper, with ± indicating standard deviation. They are not values showing how much ability fell before and after AI was introduced; they compare cases where the advice differed in correctness. The commentary's "about 80% to under 20%" corresponds to the less experienced group, and reading it as a result for all 27 readers would misidentify the subjects.
Dratsch told the RSNA:
"We anticipated that inaccurate AI predictions would influence the decisions made by radiologists in our study, particularly those with less experience."
He went on to say he was surprised that experienced physicians were also adversely affected, though to a lesser degree. Automation bias, the tendency to over-trust suggestions from an automated system, is a problem to consider even in workflows where experts do the checking.
The study did not prove decline in judgment
The 2023 experiment had no control condition in which readers worked without AI advice. It also did not compare against showing the same advice as a human opinion. It was an intervention that deliberately supplied wrong suggestions, but it cannot isolate the pure effect of using AI, or any causal effect specific to believing "the AI said so."
The higher accuracy of the most experienced group likewise reflects a comparison of existing groups, not the result of an intervention that increased experience. A short-term task with a small sample cannot support the causal claim that AI erodes human judgment in general over the long term. The commentary's concern about the future and what this experiment measured, namely reactions to wrong advice, need to be kept apart.
Follow-up work is under way. Filippo Pesapane and colleagues reported a simulated reading experiment in which six breast radiologists at one site read 200 cases, in a paper published on May 29, 2026 in the peer-reviewed European Radiology (DOI: 10.1007/s00330-026-12666-6). According to the public abstract, the study ran from March to June 2024 and compared unaided reading, AI-assisted reading, and AI-assisted reading with heat-map explanations. It reports that adding explanations did not eliminate the bias completely.
This later study also measured human reading, and it is not a direct replication that reproduced the 2023 figures under the same conditions. Because the design described in the public abstract differs from the 2023 experiment, the effect sizes of the two studies cannot be compared simply.
The RSNA summary listed displaying the system's confidence level, explaining the basis for a judgment, and users taking responsibility for their own decisions as candidate countermeasures. The 2023 experiment did not, however, test whether those measures work. Open questions remain, including how well people catch wrong advice in real deployments and whether they can retain their unaided judgment.
In assessing AI risk, it is worth checking concretely both designs that restrict access to facilities and designs that let people scrutinize outputs. If a procedure puts a human in charge of final approval, testing whether that person can spot errors under realistic conditions would provide material for judging its safety.
