On August 18, 2026, Anthropic released a technical report on de novo protein binder design using Claude, along with all designs, prompts, and measurement data. Across 15 targets for which binding could be assessed, 354 of 1,320 sequences bound, with at least one binder found for 14 targets. What was novel here was not an improvement to any individual protein foundation model. Rather, it was that a single agent—operating under a protocol set by experts—handled everything from target research to selection of publicly available tools, candidate narrowing, and ranking of sequences to hand off for synthesis.

AD

A roughly 16,000-word protocol drove the design pipeline

Claude researched targets, selected target constructs and epitopes, and then deployed specialized tools to generate and optimize candidates. From a large pool of candidates, it selected 30 designs per target, ranked them, and returned sequences for synthesis. Design decisions for each target were left to Claude, and humans did not intervene in those decisions.

However, this does not mean humans were absent from the experiment's starting point. Researchers selected the targets and assay antigens, and prepared a protocol of roughly 16,000 words (about 30,000 tokens) along with reference materials. The protocol breakdown was 34.2% science and tools, 34.7% orchestration and validation, and 31.1% operations. The latter two categories were revised during pilot campaigns and were fixed before the reported runs began. Researchers were responsible for preparing accounts and credentials, approving non-scientific access requests, recovering sessions after infrastructure failures, placing synthesis orders, and interpreting binding data.

For this reason, reading these results as Claude having created a new protein foundation model would miss the point. Claude combined already-published specialized models. Among the designs that proceeded to synthesis, ten different structure generators were used, including PXDesign (358 designs), RFdiffusion3 (267), and Genie 3 (185). SolubleMPNN was used for the majority of tested sequences (1,133), and an ensemble of ESMFold2, ESMFold2-Fast, and Protenix v2 served as the primary ranking method. AlphaFold 3 weights, Rosetta/PyRosetta, and ESM3 were excluded for licensing reasons; the pipeline was built exclusively from tools confirmed in advance to be open-source and deployable.

Recent advances in protein generation, sequence design, and structure prediction have shrunk the scale of experimental testing—from tens of thousands to hundreds of thousands of candidates in early campaigns down to tens or hundreds of candidates with newer methods. What remained was the specialized operational work: reading target biology, selecting tools, narrowing candidates, allocating compute resources, and ranking. This release makes that pipeline explicit as a lengthy protocol external to the model itself, and had an agent execute it.

354 binders, with performance split sharply by target

The released data covers 16 targets and 1,440 designs. Because measurements for mature GDF-8 were inconclusive, the analyzed set narrowed to 15 targets and 1,320 designs. Of these, 354 designs (26.8%) bound, with hits found across 14 targets. Looking only at the top-ranked design per target for each campaign, 49% bound—indicating that the ranking signal was reflected in the measurement results as well.

Run Format Binders / Designs Tested Binding Rate
Mythos Preview, 48h, multi-target 104/390 26.7%
Opus 4.8, 48h, multi-target 88/390 22.6%
Mythos Preview, 24h, single-target 158/450 35.1%

Even restricted to just the 13 targets common across formats, the single-target Mythos Preview run achieved 143/390, or 36.7%. However, single-target campaigns used 24 hours and $10,000 USD of cloud GPU compute each, while multi-target campaigns used 48 hours and $50,000 USD each. For the 13 common targets, the single-target format's per-target compute budget was 2.8 times larger. The 35.1% figure cannot be attributed to target focus alone, and the report does not separate the effects of focus and budget.

Performance varied widely by target. Pooled hit counts were 72/90 for TREM2, 54/90 for VEGF-A, and 49/90 for IL-7Rα, while BBF-14 yielded only 3/90, 15-PGDH just 1/30, and MBP 0/90. Co-folding confidence scores did not flag these latter failures in advance. For TNFα, the Opus 4.8 campaign produced 12 binders derived from four distinct backbones, while Mythos Preview produced none. These target-dependent differences show that overall average figures alone cannot capture the reliability of the design pipeline.

For RBX1, 28 binders emerged from Claude's 90 designs, compared to 9 out of 245 entries in a de novo design competition. Claude's strongest binder had a KD of 3.9 nM, versus 45 nM for the competition-winning design when resynthesized and measured on the same plate. However, of the six competition results used for this comparison, four were already present in Claude's training corpus. The commonly cited 10–15% campaign benchmark is also not a human control matched for target or budget, so this does not support a conclusion that experts were outperformed.

AD

Two CROs measured separately, but only binding was demonstrated

Measurements were conducted by two independent contract research organizations (CROs): Adaptyv Bio and Twist Bioscience. Both synthesized the sequences without modification and measured binding using different molecular formats and assay conditions. Neither party saw the other's data, nor the model, campaign, or ranking corresponding to each sequence during measurement; Anthropic then integrated the separate determinations using rules fixed in advance. This design suppresses pathways through which computational rankings or one set of experimental results could influence the other measurement.

The released data includes provenance for each design, frozen prompts, and the computational models used. On the measurement side, sensorgrams, raw data fits, and integrated binding determinations were all released. Data and documentation are under CC BY 4.0, and scripts under the MIT license; the primary dataset expands to 129,003 files and 9.9 GB. An optional tier including structural data from all predictors totals 74.5 GB. External researchers can trace not only the conclusions but also which tools and rankings each candidate passed through. That said, this is a technical report compiled by Anthropic itself, not a peer-reviewed journal paper. Re-analysis of the data and reproduction of the pipeline by other teams are separate matters that need to be considered independently.

Meanwhile, what the measurements confirmed was binding—not the three-dimensional structure or biological function of the binders. All displayed binding poses are predictions; neither the designs nor the complexes were experimentally structurally resolved, and no activity assays were performed. Affinity values for the five oligomeric targets are apparent values. Of 233 binders also tested against mouse orthologs, 130 bound, but cross-species reactivity was a secondary prompt objective.

There are also caveats regarding statistical independence. The 26.8% per-sequence figure includes variants of the same backbone; counting only the top-ranked sequence per backbone across 809 backbones yields 200 hits, or 24.7%. Each combination of model, run format, and target was run only once (except for a supplementary TNFα order), so differences due to model, format, or run-to-run variability cannot be separated. Furthermore, there is no control campaign in which human experts used the same tools and budget. The report does not claim that Claude's designs outperformed those of experts working under matched conditions.

What does "any lab" actually require?

Anthropic released data spanning design through measurement, including provenance for each sequence and the prompts used. The fact that only publicly available specialized tools were used also broadens the entry point for verifying the design pipeline or refining the procedure. But this does not provide grounds for asserting that "any lab could immediately execute this." Reproduction requires access to the models used in the report and $10,000–$50,000 USD in cloud GPU compute per campaign. It also requires wet-lab capability for synthesis and binding assays, plus the expert-written protocol. Wet-lab costs were not disclosed.

Model accessibility is also a condition. Anthropic states that it is restricting life-science research on its most capable model, Fable 5, until a trusted-access program is established, while Opus 5 will remain generally available. This campaign was run using Mythos Preview and Opus 4.8. Therefore, obtaining the public prompts and open tools does not guarantee that the same model configuration can be immediately assembled.

The next stage of verification begins with other researchers re-analyzing the released data. Beyond that, the same pipeline needs to be run under different model access conditions, compute budgets, and targets, and measurements need to extend beyond binding to structure and function. What reproduction ultimately tests, then, is not just the 354/1,320 figure, but the broader question of how thoroughly expert procedures must be codified for computational design to be run reliably and consistently.