On September 30, Amazon Web Services (AWS) released the Dogwood Local Engine under the Apache 2.0 license. It allows or denies an AI agent's tool operations by checking them against the agent's execution history. Embedded in an agent's runtime as a Rust library, it can evaluate a condition just before an operation runs, such as "allow a code push only if tests passed within the last 15 minutes and none have failed since."

The engine adds a local execution component to Dogwood, the policy language AWS released in August, and it can carry its decision state across crashes. However, the engine only decides whether to allow or deny. Actually stopping an operation in response to that decision, and accurately recording what was executed, remain the responsibility of the software that embeds it.

AD

Allow a push only if tests passed within 15 minutes and none have failed since

Dogwood is a policy language that can decide whether to allow or deny an action based not only on the current operation but also on what happened before it. When AWS announced it on August 6, it also added support to the policy feature of Amazon Bedrock AgentCore. While remaining compatible with the authorization language Cedar, it adds conditions such as the order of operations and the count or cumulative total within a given period.

The newly released Local Engine evaluates policies each time it receives an event, then saves to disk the state needed for later decisions.

The earlier Dogwood reference interpreter was meant for checking language specifications and sample behavior, not for direct use in production authorization. The new library can process a continuous stream of operations and keep making decisions that take past state into account, even after a restart.

The "Local" in the name means the engine can be embedded in the user's own runtime. It does not mean the LLM itself runs locally.

In the code-development example AWS presented, running tests is always allowed. A Git push, which sends changes to a remote repository, requires that tests passed within 15 minutes and that no test has failed since.

A push right after a test failure is denied, and a push 5 seconds after a fixed test passes is allowed. Once 17 minutes have passed since the last success, it is denied again.

The history consists of two kinds of events: a request to call a tool, and the result of that execution. The engine decides to allow or deny on the request event, while the result event is added to the history as input for later decisions.

That means a request to run tests alone does not satisfy the condition. Only when the success or failure result returned by the test tool is also recorded can a later push be judged correctly.

Separating the engine that decides from the runtime that actually blocks

The Dogwood Local Engine itself has no function for stopping a command it has denied.

The runtime that mediates tool calls passes each request to the engine and executes the tool only if it is allowed. After execution, it sends the result back to the engine. A denied request never runs the tool, so no result event occurs.

In other words, the Local Engine handles decisions and restoring state after a crash, while the host runtime is responsible for actually blocking operations and keeping any required audit records.

Organizing the official implementation notes and API documentation, the roles divide as follows.

Process Mainly handled by Conditions and limits for it to work
Allowing or denying operations Engine Required events must arrive without omission and with accurate content
Blocking denied operations Runtime Only tool calls the engine has allowed may be executed
Restoring decision state after a crash Engine The state needed for decisions must be recoverable from protected snapshots or remaining logs
Auditing past operations and decisions Runtime A complete audit record must be kept separately from the engine's internal persisted data

This table organizes the explanation for the 1.0 series, as published on October 2, 2026, into four items: decision, blocking, recovery, and audit. The premise is that required events are input accurately and that the runtime can reliably apply the decision results. It is not an evaluation measuring resistance to attacks.

When adopting it, simply making the engine's decisions available is not enough. You also need to confirm that the AI agent cannot perform the same operation through another route that bypasses Dogwood's decision.

The events themselves must also be trustworthy.

The Local Engine cannot independently verify things like who performed an operation or whether a test truly passed. Authenticating callers and accurately recording execution results must be provided by the runtime.

For an AI agent that can operate a shell, the saved files and the engine's memory also need protection, through OS access controls, sandboxing, or similar measures, so they cannot be altered via tools.

The code-development example carries another caveat. The condition AWS showed checks whether tests succeeded or failed, but it does not check whether the code that was tested is the same as the code about to be pushed.

If you want a production condition such as "only verified changes may be sent," you need a design that includes information identifying the code version, such as a commit ID, in the events, and checks that the successful test was for the same target. A time condition like "succeeded within 15 minutes" does not guarantee that.

AD

With concurrency, the difference between "when requested" and "when completed" matters

When multiple tool calls arrive at the same time, the Local Engine uses a lock to process them one at a time.

It timestamps each event, appends it to the log, finishes syncing to disk, and only then evaluates the policy. Because the lock is held until the decision is finished, a request that arrives later is evaluated against a state that includes the events recorded before it.

However, what is guaranteed here is only the order in which events used for decisions are recorded. It does not guarantee that externally executed tools finish one after another.

For example, if a push request arrives while a test is still running, then even if that test later fails, the failure result does not yet exist in the history at the moment the push is judged.

To make a rule such as "also forbid pushes while tests are running," you need to express the state "test started but not yet completed" as a condition too.

In its August Dogwood announcement, AWS explained how the choice of what to count under concurrency matters, using a transfer limit of "$5,000 per hour" as an example.

Consider three consecutive $2,000 transfer requests, none of which has completed yet.

Under a policy that sums only the results of completed transfers, the amount of "completed transfers" is still zero when the third request is judged. All three $2,000 requests would be allowed, and a total of $6,000 could be transferred.

Under a policy that sums transfer request events, the request currently being judged is also included in the tally. At the third request, the requested total reaches $6,000, so it can be denied there.

However, with this method, denied requests also remain in the history. It therefore has a different meaning from a limit on "the amount actually transferred."

Should the policy count the "attempted amount," including failures and denials, or the "transaction amount" that actually completed? Even with the same $5,000 limit, the behavior changes depending on which events are counted.

The Local Engine makes it possible to evaluate concurrently arriving events in a consistent order. But which stage of events a business rule should target must still be decided by the user.

After a rule change, past successes are not carried into the new condition

Time conditions also bring a caveat when policies are changed while the system is running.

A newly added policy, or a policy whose time condition has been changed, is in principle evaluated using events that arrive after the update. It does not retroactively bring events from before the update into the new condition's history.

On the other hand, according to the API documentation, for policies that have not changed, the decision state accumulated so far is preserved.

AWS explains this using lint, a static code check, as an example.

Suppose lint succeeds first, and then a rule is added saying "a commit is allowed only after lint has succeeded." In that case, a commit right after the rule is added is denied.

Lint has already succeeded, but that happened before the new rule took effect. If lint is run again after the policy update and succeeds, subsequent commits are allowed.

This behavior is also related to the fact that the Local Engine is not designed to keep every past event forever.

The engine gathers the information needed to evaluate current time conditions into in-memory state and periodically saves it to a snapshot. Old logs already reflected in a snapshot are pruned, so the information a newly added condition needs may not remain as a complete history of past events.

Evaluating new conditions using only events after the policy update avoids this problem.

The policy update itself is also recorded in the same log as events, which fixes its order relative to ordinary requests. This prevents a single request from being judged under a mix of the old and new policies.

On restart, policy updates are replayed in order as well.

However, if a policy update fails, the engine keeps running with the policy it was using just before. The embedding software must provide a way to notify operators of update errors so users do not mistakenly believe the change succeeded.

Being able to persist state is also a separate matter from being able to keep audit records of every operation.

According to the official implementation notes, the internal log is periodically consolidated into a snapshot and then pruned. Decision results themselves, such as allow or deny, are also not recorded for audit purposes.

In other words, the state needed to continue making decisions after a crash can be recovered, but if you want to trace everything about who tried what in the past and why it was denied, you need to prepare a separate audit log.

AD

Time window length and operation granularity affect decision speed

In AWS's performance measurements, lengthening the period that a time condition references greatly increased evaluation time.

When AWS simulated a session generating 360 events per hour, a 15-minute window limited the target to the most recent 90 events, so evaluation time plateaued at a steady level even as the session grew longer.

With a 24-hour window, by contrast, the amount of data referenced kept growing, and at the 12-hour point of the session a single evaluation took about 6 milliseconds. That was about 300 times the 15-minute case, according to AWS.

The measurements were taken on a single server-class host, using the median of 200 evaluations at each measurement point. AWS says variation between repeated runs was within 5%.

However, these figures show only the evaluation time of the policy conditions themselves.

The time to sync events to disk depends on storage performance, so the figures do not directly show the total wait an AI agent experiences for a decision. Nor can they be read as the performance of the whole system including actual tool processing.

How finely operations are divided also affects performance.

AWS compared two ways of defining operations, using a setup in which 100 policies were assigned 20 each to five kinds of Git operations.

If everything is grouped as a single git operation and push and commit are distinguished by looking at arguments, even a push request requires evaluating all 100 policies.

If operations themselves are defined separately, such as git:push and git:commit, only the 20 relevant policies need to be evaluated on a push. In this simulation, the latter was about 5 times faster.

That said, shortening the time window is not an automatic way to improve performance.

For example, if a business rule such as "limit the cumulative amount over the last 24 hours" is shortened to 15 minutes purely for performance reasons, the meaning of the rule itself changes.

Also, although time conditions are given a mathematical definition, AWS said in its August announcement that the automated reasoning analysis that Cedar offers does not support Dogwood's time conditions.

In other words, it does not automatically prove that a business rule written in Dogwood is what was intended.

When considering adoption, you need to measure wait times under real load, including disk sync, and confirm that the intended conditions work correctly even with concurrent requests and after policy changes.

If those prerequisites are met, and the runtime also reliably blocks and records operations, the Dogwood Local Engine can serve as a foundation for controlling how much tool use to entrust to AI agents, taking past execution results into account.