On October 9, Anthropic added "dynamic workflows" as a public beta to Claude Managed Agents, its service for running AI agents in the cloud. Claude writes a program that manages how work is divided and how it proceeds, runs multiple agents, and compiles their results. According to the official release notes, this lets users continue conversing with the lead agent while large-scale document reviews and code audits run in the background.

A single run can launch up to 1,000 agents in total. However, only 64 threads can run concurrently at present. That does not mean 1,000 agents work at once, nor that processing becomes 1,000 times faster.

What matters about this feature is that everything from dividing up work to re-verifying results can now proceed automatically through a program. To judge its practical value, you need to look not only at the number of problems found, but also at how many targets went uninvestigated and how much the run cost.

AD

Claude writes the program that assigns work and manages progress

Dynamic workflows were announced for the developer tool Claude Code on May 28. The latest announcement brings the mechanism to the cloud execution environment of Claude Managed Agents. Although the feature has the same name, the execution environment and limits differ, so the two should be considered separately.

In conventional delegation to subagents, the lead Claude decides what to hand off next each time, receives reports from the assigned agents, and then moves on to the next task.

With dynamic workflows, a program handles this coordination instead. It receives each agent's results, passes them to other agents, and repeats steps as needed. According to the official feature description, the lead agent can keep responding to the user and checking progress in the meantime.

For example, when reviewing a large number of contracts, you could assign one agent to each contract to check its contents, then have different agents re-check all of the contracts.

If only contracts where the first pass found problems are re-verified, any issue the first agent missed will go unnoticed. In the sample instructions Anthropic presents, therefore, a separate agent checks every contract regardless of the first-pass results. If a contract fails verification, the work is redone and verified again.

In this way, what to investigate, in what order, and under what conditions to re-verify can be built into the workflow in advance.

Using dynamic workflows does not mean all work is automatically double-checked, however. How work is divided, what is subject to verification, and how failures are reported depend on the user's instructions and the workflow design.

To use the feature, set multiagent.type in the agent definition to multiagent_20261001 to enable workflows. With this setting, workflows and delegation to subagents are both enabled by default.

After that, simply describe the work you want done in a message or system instruction, and the lead agent decides whether to start a workflow. No separate API call is needed to launch one.

Up to 1,000 agents can be launched, but only 64 threads run at once

The workflow run specifications set separate limits on the cumulative number of agents that can be launched in one run and the number of threads that can be active at the same time.

A thread here means a unit in which each agent works with its own conversation history. The main limits as confirmed on October 10, 2026 are as follows.

Limit Current cap / setting Notes
Cumulative agents launched per run Up to 1,000 A cumulative total across the whole run; it does not mean 1,000 agents run at once
Concurrent threads per run 64 The next agent cannot start until a slot opens. This value may change and is not an API-guaranteed figure
Run lifetime 24 hours by default Can be shortened on the agent side. A run may also expire while waiting
Unfinished runs retained per session 10 by default Includes runs paused because they hit a budget cap, for example

When an agent fails, the server side may re-run the work in a new thread. As a result, the cumulative number of threads used across a run can exceed 1,000, and the thread count shown on the monitoring screen does not always match the number of agents launched.

Separate per-organization rate limits on model usage also apply. Even if you can launch agents up to the cap, there is no guarantee that processing will complete at a given speed.

Each agent's conversation history is independent, but the files being worked on are shared in the same execution environment.

This spares the lead agent from having to load large volumes of findings directly, but it does not prevent conflicts when multiple agents rewrite the same file. When using the feature for tasks such as code changes, you need to clearly separate which files and what scope each agent is responsible for.

The agents that receive work can be registered in advance or defined anew at run time. Agents defined at run time use the same model as the lead agent.

If you want to hand some tasks to a different model, you need to register an agent that uses that model beforehand. For registered agents, the tools and MCP servers they use can be specified individually, which makes it easier to limit what each one can access according to its job.

Note that the lead agent cannot send additional messages later to an agent running inside a workflow.

This differs from ordinary subagents, with which you can continue the conversation with the same agent after receiving its report. If you need repeated corrections or re-verification, that processing must be built into the workflow program.

AD

66 of 70 bugs found, a large gap from a single agent

Anthropic published the results of an internal test on its official developer account, in which 70 bugs were deliberately planted in 116,000 lines of code.

When it ran the investigation three times each with a single agent and with a dynamic workflow, the single agent found 14, 15, and 27 bugs. The dynamic workflow, by contrast, reportedly found 66 in all three runs.

Against the 70 bugs, the single agent's detection rate was 20.0% to 38.6%, while the dynamic workflow reached 94.3% in all three runs.

Run Bugs found (single agent) Bugs found (dynamic workflow) Detection rate (single / workflow)
1st 14 66 20.0% / 94.3%
2nd 15 66 21.4% / 94.3%
3rd 27 66 38.6% / 94.3%

Source: Results of an internal test published by Anthropic on October 9. The detection rate is the number found in each run divided by the 70 planted bugs, multiplied by 100 and rounded to one decimal place. It does not represent accuracy across all problems in the code, nor a false-positive rate.

The results suggest that dividing work among multiple agents and verifying it may reduce oversights when investigating large codebases.

Still, even the dynamic workflow missed four bugs every time. And the figure of 66 found in all three runs does not by itself show whether the same 66 were found each time. Nor does it show that the bugs found were fixed correctly.

Furthermore, the published detection counts alone do not reveal whether both approaches were compared on the same token budget, or how many false positives occurred.

Efficiency, including cost and processing time, also cannot be assessed from these numbers. Whether similar results can be obtained with codebases of different sizes and types will need further verification.

More agents mean more cost: watch budget controls and completion status

Dynamic workflows themselves carry no separate usage fee. Charges accrue for the tokens each agent consumes, at the rates of the model used.

In addition, Managed Agents charges $0.08 per hour of session runtime.

Even when multiple agents run in parallel within the same session, overlapping runtime is not counted twice. Running 64 agents at once therefore does not multiply the hourly rate by 64.

On the other hand, every time an agent reads materials, reasons, and verifies results, token consumption rises accordingly.

In its January 23 explanation of designing multi-agent systems, Anthropic said implementations commonly consume 3 to 10 times more tokens than a single agent handling equivalent work.

This is not a measurement of consumption by the new dynamic workflows, but it indicates that even if parallelization shortens the time to completion, total token consumption and cost do not necessarily fall.

To manage such costs, Managed Agents lets you set a budget cap when creating a session.

The budget is shared by all agents in the workflow. The cap is judged not on the actual billed amount, which reflects contractual discounts and the like, but on usage calculated from list prices.

Once the cap is reached, new model requests stop being issued. However, requests already in progress run to completion, so the set budget can be exceeded by up to one request per active thread.

Reaching the budget cap also does not mean the workflow has finished.

The run enters a paused state and can be resumed by raising or removing the cap. But because the run's lifetime keeps elapsing while it is paused, it may expire before it is resumed. When designing long-running processes, you need to account not only for the budget but also for the time spent waiting or paused.

More important still, a run result showing completed does not necessarily mean all the work finished successfully.

The official specifications state explicitly that the workflow program itself may complete even if some agents fail at their work or required agents cannot be launched.

For example, when reviewing a large number of contracts, you need to have the workflow report not only the results for contracts that were checked but also, without omission, those that could not be read or failed verification. For code audits, you should have it state the scope that went uninvestigated, in addition to whether the problems it flags can be reproduced.

Anthropic recommends dividing work into units that each agent can process independently.

If you split development of a feature by stage into "planning," "implementation," and "testing" agents, design intent and work status have to be handed over repeatedly. It is easier to take advantage of this feature by assigning work per independent document or per piece of code with a clearly defined scope, and building the necessary verification into the workflow.

When actually adopting it, it is advisable to start by narrowing the target and comparing against a single agent using the same materials and evaluation criteria. You should then record the scope investigated, the accuracy of the problems found, the number of oversights and unprocessed items, and the cost.

If it can investigate all the necessary targets and deliver reproducible findings within budget, dynamic workflows could be a strong option for streamlining large-scale document reviews and code audits. How far AI can be trusted with steps in which humans previously directed each piece of work will be the focus of its practical adoption going forward.