An implementation report on "AI-TEC" (AI-Agent Augmented Tsinghua Eye Clinic), an AI-agent ophthalmology clinic run by a Tsinghua University research team at Beijing Tsinghua Changgung Hospital, was published in Nature Medicine on September 10, 2026. The report states that diagnostic accuracy for fundus images improved after adding a small number of high-quality images that specialists had relabeled to the training data. Even so, in the one-month period five months after launch, only 41 of 1,113 total exams went through the AI-TEC workflow. The following month, that number rose to 259 out of 1,126 exams. The uptick in usage came right after the on-screen operating procedure was revised.

AD

An eye clinic redesigned around AI, and labels redone by specialists

The design philosophy behind AI-TEC does not treat AI tools as something bolted onto existing clinic operations. Instead, everything from pre-visit intake to post-visit follow-up is restructured on the assumption that an AI agent is present throughout. ScienceAlert's report describes the framework this way. Ophthalmologists are not removed from the process—neither the publication page nor the news coverage contains any indication that AI alone determines diagnosis and treatment. The lead authors are Tao Yan and Di Zhang, with four corresponding authors: Jiamin Wu, Ya Xing Wang, Qionghai Dai, and Tien Yin Wong. Their affiliations span Tsinghua University's Department of Automation, Tsinghua University School of Medicine, the Eye Center at Beijing Tsinghua Changgung Hospital, and the Singapore Eye Research Institute.

What's reported about how the training data was handled runs counter to the usual machine-learning intuition. The initial dataset consisted of roughly 27,000 fundus images that were low-quality and incompletely labeled. Retraining with a small number of high-quality images that ophthalmology specialists had relabeled correctly sharply reduced false positives and improved diagnostic accuracy for glaucoma and age-related macular degeneration, according to ScienceAlert. Rather than adding more data volume, it was adding a handful of images vetted by specialist judgment that made the difference.

However, accounts of exactly how many images were added don't agree: ScienceAlert reports 1,426 images, while another tally puts the figure at 1,482. The specific accuracy values reached for glaucoma and age-related macular degeneration are similarly inconsistent across reports. Rather than listing numbers, what remains solid is this: it wasn't a new model architecture or more data that boosted accuracy, but the manual process of specialist labeling.

The research team states, in effect, that implementing AI-centered medicine is fundamentally an ecosystem challenge, and that its effectiveness depends not on individual algorithm performance but on data quality, workflow interactions, clinician engagement, and measurable clinical value (as quoted by ScienceAlert). This statement should be read not merely as a summary of the implementation report, but as a setup for the usage-rate figures that follow.

After accuracy improved, month five's tally was 41 cases

In the one-month period five months after launch, 1,113 exams were conducted at the ophthalmology department of Beijing Tsinghua Changgung Hospital. Of these, only 41 went through the AI-TEC workflow. The following month, after modifications were made to speed up the process by reducing clicks and manual data entry, 259 out of 1,126 exams went through the same workflow.

Calculated from the reported figures, the proportion of exams that went through the AI workflow rose from 3.7% to 23.0% in a single month—a 6.2-fold increase. This jump occurred right after the operating procedure was revised. The math is straightforward: 41 divided by 1,113 is 3.6837%, and 259 divided by 1,126 is 23.0018%. Dividing the unrounded figures gives a 6.24-fold increase, which rounds to 6.2 at one decimal place. Both denominators are the same facility's total monthly exam count, with units aligned as "exams."

Two caveats apply to this calculation. First, there's a discrepancy in the source: ScienceAlert states the previous month's rate as 3.8%, but the same article's figures of 41 and 1,113 only yield 3.7%. It seems more sensible to prioritize the raw counts and use the calculated value.

Second, there's the question of causal scope. It's not certain that the operating procedure was the only thing that changed between these two consecutive months—no material has been reported that rules out seasonality or differences in patient composition. The exam counts themselves were obtained via ScienceAlert.

Even so, the gap between these two months bears directly on how medical AI should be evaluated. Raising diagnostic accuracy and raising the frequency of actual clinical use are different jobs requiring different skills. The former becomes a published paper; the latter gets treated as an unremarkable interface tweak. The Nature Medicine publication page itself states, in its title, essentially that moving from an AI-assisted tool to an AI-centered model of care requires workflow integration, clinician engagement, and measurable clinical value.

AD

AI reasons from outcomes; doctors reason from symptoms

One remaining issue reported is a mismatch in the order of reasoning. AI starts from outcomes. Given a fundus image, it answers, with a probability, whether this patient has glaucoma or age-related macular degeneration. Ophthalmologists, on the other hand, start from symptoms. When a patient comes in complaining of blurred vision, the doctor works to narrow down the cause.

This difference generates friction in the clinic, and tracing it step by step makes the mechanism clear. The question the doctor is asking is "What is causing this blurred vision?"—with candidates for differential diagnosis ranging from cataracts to refractive error to retinal disease. But what the AI returns is an answer about the presence or absence of specific diseases, and that answer to a question the doctor hasn't yet posed appears on screen first. The doctor then has to re-translate it into their own differential list and personally reconcile it with the symptoms.

The translation effort is small per case, but clinic visits have fixed time allotments per patient. As clicks and manual entries accumulate, the judgment tips toward skipping the step entirely as the faster option. The report cites this kind of burden as the background for the drop to 41 cases at month five.

The same problem framing has also been raised within ophthalmology itself. In a commentary published in Eye on May 20, 2026, Devina Ramesh, Brendan K. Tao, and Cindy M. L. Hutnik identified barriers to adoption in the fact that conventional AI models were built for single tasks, couldn't perform longitudinal reasoning, and were integrated into workflows in only a one-dimensional way. The three authors positioned AI agents that bundle multiple models and can detect discrepancies among outputs as a solution to this. Physicians remain in the position of deciding whether to accept or reject each individual suggestion. A research group overlapping with the AI-TEC authors also laid out the same problem—moving from model to clinical workflow—in a Viewpoint published in JAMA Ophthalmology on July 2, 2026.

Risks in the opposite direction have also been reported. In a study of four endoscopy centers in Poland by Budzyń et al. (The Lancet Gastroenterology & Hepatology, 2025), endoscopists who had gained experience with AI assistance showed lower adenoma detection rates when performing standard colonoscopies without AI. The rate was 28.4% (226 of 795 cases) before AI adoption, versus 22.4% (145 of 648 cases) after.

The drop in performance occurred specifically in exams conducted without AI. There are problems with a design that goes unused, and problems remain even after a design becomes overused. Workflow adjustment has to address both at once.

Who built it, and who funded it

The Nature Medicine publication page includes a conflict-of-interest disclosure. One of the corresponding authors, Tien Yin Wong, is listed as a co-founder and patent holder at EyRiS, a startup developing digital solutions for eye disease. He is also listed as a consultant for companies including AbbVie, Bayer, Boehringer-Ingelheim, Carl Zeiss, Genentech, Novartis, Roche, and Sanofi. The other authors declare no conflicts of interest.

That disclosure, however, is absent from the news coverage that widely reported on this study. One of the paper's corresponding authors discloses that he is a co-founder of a company developing digital products for eye disease, yet ScienceAlert's report—the starting point for this article—does not mention that disclosure. The method of comparison is simple: read the Ethics declarations section of the publication page and the body of the news report directly, and compare them on a single point—whether conflicts of interest and funding sources are mentioned. That said, only ScienceAlert's full article was verified; the full text of other outlets that reconstructed the same content was not checked.

The existence of the disclosure itself is not a problem. If anything, it's natural for a researcher who has worked in the ophthalmology AI field for a long time to have ties to related companies when participating in a clinical implementation report. There is no mention, in either the publication page or the news coverage, that EyRiS was involved in AI-TEC.

Still, readers need information about who was involved in the design and who funded it in order to calibrate their own assessment of an AI-centered clinic framework. As for funding sources, the publication page explicitly lists the National Natural Science Foundation of China, two key laboratories under the Beijing municipal government, and the XPLORER PRIZE from the New Cornerstone Science Foundation. The fact that this research was supported by Chinese public research funding and foundations is not unrelated to the national policy story discussed below.

AD

China put institutions in place first; Japan is still just building models

The AI-TEC report is not a one-off research result. It emerged after two years in which pricing categories, institutional bodies, and national policy statements were assembled in sequence. Laying them out in order makes the trajectory visible.

In November 2024, China's National Healthcare Security Administration created an "AI assistance" expansion item for pricing categories covering radiology, ultrasound, and rehabilitation services. At the same time, it stated that patients could not be separately charged for the AI-assisted portion on top of existing fees. On April 26, 2025, Tsinghua University held a launch ceremony for its AI Hospital, announcing plans for trial operation at Beijing Tsinghua Changgung Hospital and the university's internet hospital. Ophthalmology was one of the pilot sites included.

On October 20, 2025, five government departments including the National Health Commission published implementation opinions on the application and development of "AI + healthcare," laying out 24 priority applications across eight directions along with two-stage targets for 2027 and 2030. Then, on September 10, 2026, the AI-TEC implementation report was published.

This sequence is a juxtaposition of dates, not proof of causation. The pricing category set by the Healthcare Security Administration is sourced from Chinese media commentary, and the original text of the notice itself could not be verified within publicly available sources. Whether AI-TEC is the same project as the ophthalmology pilot under Tsinghua's AI Hospital cannot be determined from the information published. Still, the sequence—in which pricing structures, institutions, and national targets were put in place before a clinical implementation report emerged—is worth reading as background for the research.

Lining up the three countries by payment system makes the differences starker. The United States grants independent reimbursement for autonomous AI diagnosis of diabetic retinopathy; China treats AI assistance in radiology, ultrasound, and rehabilitation as an expansion of existing fee items without allowing separate billing; and Japan is still at the budget-request stage for a project implementing agent-based AI to support clinical operations.

In the United States, IDx-DR (now LumineticsCore) received FDA De Novo authorization in 2018 as an autonomous AI diagnostic tool for diabetic retinopathy. The clinical trial that underpinned the approval, by Abràmoff et al. (npj Digital Medicine, 2018), reported 87.2% sensitivity and 90.7% specificity for detecting moderate or worse diabetic retinopathy. A reimbursement code was subsequently established for automated retinal imaging, creating a pathway through which AI-based judgments earn independent compensation.

China classifies AI assistance in radiology, ultrasound, and rehabilitation as an extension of existing pricing items and does not, within that scope, allow additional billing for the AI component alone. Japan already has an AI medical device covered by insurance since December 2022, such as nodoca for influenza diagnosis support, but for business support using agent-based AI, the Ministry of Health, Labour and Welfare has only newly included a "Development and Implementation Project for Business Support Using Agent-Based AI in the Medical Field" in its fiscal year 2027 budget request, with a reported requested amount of 2.5 billion yen.

The axis of comparison here is narrowed to a single point: how public payment systems treat AI-based diagnostic support. Regulatory approval and insurance reimbursement are separate systems, so approval-related matters are not mixed in here. Because the three countries differ in target diseases and scope of application, the monetary figures themselves cannot be directly compared. Japan's requested amount is a figure reported in commentary articles.

Even so, the difference in institutional positioning is discernible. In a country where AI can earn independent compensation, there's also a revenue-based reason to use it in the clinic. In a country where using it doesn't change the billed amount, the only reasons to use it are speed and quality of operations—and a high click count directly becomes a brake. Japan is still at the stage before that, deciding what to build.

The numbers to check first once the original paper is public

A number that will shape how this implementation report should be read remains unconfirmed. How many high-quality images did specialists relabel? What accuracy did the model reach for glaucoma and age-related macular degeneration? The publicly available scope of the publication page and the descriptions in secondary reporting diverge on exactly these two points. Once the full text of the paper becomes readable, these are the first things to cross-check.

The publication page has no registered abstract, lists 15 references, and carries only one publication type in PubMed—"Journal Article." Whether this is peer-reviewed original research or an editorial-style contribution would also affect how the figures of 41 and 259 should be understood.

Gaps remain on the operational side as well. When AI-TEC began operating and how many patients it has seen overall have not been disclosed. Even the five-month starting point is only a relative timeframe indicated by the report. Whether the 259 figure held up after the interface was simplified, or fell again afterward, will have to wait for further reporting. On the Japanese side, tracing the Ministry of Health, Labour and Welfare's published fiscal year 2027 budget request documents would allow further confirmation of what the agent-based AI business support project targets and what evaluation metrics it sets.

What the research team itself puts forward as its conclusion is the view that algorithm performance is not what determines the success of implementation. What matters, they say, is data quality, workflow interactions, clinician engagement, and measurable clinical value. If this view is correct, what's needed next is not a new model. It's a system for continuously running specialist labeling, an interface that returns output in the order matching a doctor's differential reasoning, and a payment structure that returns compensation to the facility for actual use. Once all three are in place, the AI agent becomes a step that actually gets used during the wait time in the clinic.