AWS Open Source Blog
Introducing the Dogwood Local Engine: temporal governance for agent actions
In August, we released Dogwood, an open source governance language for agents and their actions. Dogwood policies specify which actions an agent may take and under what conditions. A key feature of the language is support for temporal conditions, which let a policy refer to an agent’s past actions and their outcomes. A Dogwood policy can require that certain past actions have or have not occurred, impose constraints on the outcomes of those prior actions, and also restrict their order and timing. For instance, Dogwood lets you write a policy that permits a coding agent to perform a git push only if a repository’s tests were successfully run within the last fifteen minutes and no test run has failed since. To enforce such policies, we need an engine that decides, for each action, whether the policies permit it.
Today, we’re releasing the Dogwood Local Engine under the Apache 2.0 open source license. The engine is local in the sense that it can be embedded as a library into an enforcement layer that regulates agents’ tool calls. To allow or deny actions based on temporal conditions, the engine maintains a record of past actions, tracking the order in which the actions occurred, and also their timing and outcomes. Since the engine’s decisions depend on this record, the engine must maintain it carefully, even in the presence of concurrent actions. In addition, the engine must keep track of prior actions in a durable way, so that even if the system crashes, it can account for them after restarting.
The rest of this post describes how the engine works and how it can be used to regulate agents. We start with an example, explaining how to write a Dogwood policy, and then illustrate how this policy would judge a sequence of actions. We then describe how the engine should be integrated with an agentic harness, the component that runs agents and mediates tool calls, and what the harness must do with the allow/deny verdicts issued by the engine. Next, we describe how the engine maintains a consistent view of a sequence of concurrent actions, even in the presence of crashes. We then show how the engine supports dynamic policy updates, letting users update their policies mid-session without pausing their workflow. Finally, we measure the engine’s evaluation time for a request and discuss schema and policy design decisions that influence the latency.
Example Dogwood policies: run tests before pushing
Let’s start by considering Dogwood policies that permit a coding agent to push only when its most recent test run passed, and that run finished within the last fifteen minutes. Here is how you can state that rule in Dogwood:
// Running the test suite is always permitted.
@id("run_tests")
permit (principal, action == Sandbox::Action::"tests", resource);
// Permit a push only if a test run passed in the last 15 minutes
// and no run has failed since.
@id("push_after_green_tests")
permit (principal, action == Sandbox::Action::"git:push", resource)
when temporal {
!Sandbox::Action::"tests"::response{ output.passed: false }
since within 15m
Sandbox::Action::"tests"::response{ output.passed: true }
};
Each permit statement above is a policy. The policy run_tests says the agent may always run the tests. The policy push_after_green_tests says the agent may push only when a test run has passed within the last fifteen minutes and no run has failed since. The when temporal clause expresses that condition. The clause is temporal because whether it holds depends on the events that happened before the request.
An event is a timestamped record of one step of a tool call. A tool call is typically recorded as two events: a request event when the agent calls the tool and a response event when the tool returns its result. The events form the history against which the engine evaluates temporal conditions. The push_after_green_tests policy judges based on tests::response events, capturing the result of the test tool. The engine only issues allow/deny verdicts for request events. Response events record the outcome of the call, such as whether the tests passed, enabling temporal conditions to judge based on those outcomes.
To better understand how these policies behave, here is one example session represented as a list of events. Each line records an event. The event description on a line starts with the timestamp “@n” indicating that the event started n seconds after the system began recording history, then it lists the kind of event, followed by the arguments and results associated with the event. For request events, at the end of the line, we list the verdict that the local engine would give. Here, we will just list the tests and git:push events that are relevant to the above policies. We omit the principal and resource associated with each event for brevity:
@10 tests::request { suite: "unit" } -> ALLOW // run_tests: tests are always permitted
@55 tests::response { suite: "unit", passed: false }
@60 git:push::request { ref: "refs/heads/fix-latency" } -> DENY // push_after_green_tests: latest run failed
@100 tests::request { suite: "unit" } -> ALLOW // run_tests: tests are always permitted
@145 tests::response { suite: "unit", passed: true }
@150 git:push::request { ref: "refs/heads/fix-latency" } -> ALLOW // push_after_green_tests: green run within 15m, none failed since
@1200 git:push::request { ref: "refs/heads/fix-latency" } -> DENY // push_after_green_tests: green run aged out of the 15m window
At @10, the agent calls the test tool, and the engine allows the request. At @55, the tests finish with a failure, and the engine records the response. At @60, the agent tries to push. The only test run in the history is a failure, so the engine denies the push. Let’s imagine that the agent fixes the code and runs the tests again at @100, and this time the tests pass at @145. When the agent pushes at @150, the passing run is five seconds behind it, and the engine allows the push. Then the agent moves on to another task. By the time it pushes again at @1200, seventeen minutes have gone by. The passing run is no longer within the last fifteen minutes, and the engine denies the push again.
The policy and the events share a vocabulary, such as the action tests and the field output.passed. That vocabulary is declared in a schema. For each action, the schema defines the principals that can take it, the resources it can act on, and the fields in the request’s input and the response’s output:
namespace Sandbox {
entity Agent;
entity Box;
type TestInput = { suite: String };
type TestOutput = { passed: Bool };
type PushInput = { ref: String };
action "tests" appliesTo {
principal: [Agent],
resource: [Box],
context: { input: TestInput, output?: TestOutput }
};
action "git:push" appliesTo {
principal: [Agent],
resource: [Box],
context: { input: PushInput }
};
}
Here, a tests response carries output.passed as a boolean, which is the field used by the temporal clause to check the outcome of the run. The schema lets Dogwood catch mistakes before a policy is deployed. Its validator checks each policy against the schema and rejects a condition that uses a field the action does not define.
Verdict enforcement and event integrity
The engine is a library that observes a stream of events and returns a verdict for each request. However, it does not directly enforce the verdict. Some other software component is responsible for acting as an enforcement layer that intercepts agents’ tool call requests, sends the details to the local engine, and then either runs or blocks the tool call based on the verdict. We’ll call this component the agent’s harness. In practice, an implementation might directly integrate the local engine by modifying an existing harness software framework or by adding an additional sandboxing layer around a harness. For simplicity, we’ll use harness as an all-encompassing term for this enforcement layer in the rest of this post.
In order for the local engine to work correctly and ensure the intended policy behavior, the harness must provide two properties:
- Verdict enforcement: when the engine denies an action request, the harness must ensure the action does not run.
- Event integrity: the engine must receive events only from the harness, with fields that accurately describe the event.
To accurately handle verdict enforcement, the harness intercepts every tool call, submits a request event to the engine, and runs the tool only if the engine’s verdict is allow. After the tool returns, the harness submits the response event. The diagram below shows one allowed call and one denied call. Note that in case of a denied request, there is no response event, because the tool didn’t run.
For event integrity, the event interface must be reachable only from the harness. Moreover, it is important that the harness and engine’s state cannot be manipulated by other processes running on the system, or by the effects of tool calls. For example, when using the harness and engine to regulate an agent’s ability to execute shell commands and run processes, the harness must be designed to use operating system isolation and sandboxing primitives so that the agent cannot invoke a process that tampers with the engine’s state. The appropriate mechanism for isolation will depend upon the specifics of the system on which the harness is running. The Dogwood Local Engine library itself does not directly perform this kind of isolation. Its responsibility is solely to issue correct allow/deny verdicts based on the events it is shown and the set of policies.
Crash-safe evaluation of concurrent requests
A temporal policy judges a request based on the sequence of events that came before. As described in the previous section, the engine relies on the harness to send it all of the events. Once the events reach the engine, it is the engine’s responsibility to correctly manage its internal state and properly track the events it has seen. This requires addressing several challenges. First, temporal policies operate over a totally ordered sequence of events in the history, yet the engine may receive multiple events concurrently, such as when one agent calls tools in parallel or several agents share the engine, and concurrent submissions have no natural order. The engine must therefore decide an order and evaluate requests against it. Second, a restart or a crash ends the process and its working memory. If the engine tracked prior events only in memory, it would lose events at restarts and assess later requests against an incomplete history. The engine must therefore make its state durable by storing it on disk and resuming from this state after a restart.
The engine processes an event in three steps: it linearizes the event, persists it, and then evaluates the policies. To linearize concurrent submissions, the engine uses a lock that admits one submission at a time. Under the lock, the engine stamps the event with the time it arrived and appends it to the event log. The event’s position in the log decides its place in the sequence used for evaluating temporal policies. The engine syncs each append to disk before evaluating the event to ensure the event is registered in a durable way.
Then, still under the lock, the engine begins evaluating the policies. In order to evaluate temporal clauses efficiently, the engine does not directly inspect the event log. Instead, it maintains a map from each temporal policy clause to a more efficient in-memory summary that tracks parts of the history that are relevant for evaluating that clause. For example, for the push_after_green_tests policy above, it is not relevant to track details about all prior events. Instead, it suffices to track just the test runs that have occurred within the past 15 minutes and whether a failure has occurred on the latest run. On each incoming event, the engine can efficiently update the summary information needed for this temporal clause: new runs get recorded in a list and are removed when they age out of the 15 minute window.
Researchers have developed efficient algorithms for maintaining summary information needed to evaluate temporal formulas. The implementation in the local engine is inspired by some of these algorithms, although currently it does not support all of the optimizations that have been described in prior research in this area.
If the engine crashes or is shut down, the in-memory evaluation state is lost and must be rebuilt after the engine restarts by replaying the event log. To keep the log bounded, the engine periodically saves a durable snapshot of its in-memory state and prunes older events from the log that are already captured in the state snapshot. On restart, it loads the snapshot and replays only the events that arrived after the snapshot was created. We implemented the log with redb, a pure-Rust embedded key-value store, so a harness that links the engine needs only the Rust toolchain.
Dynamic policy updates
The engine judges requests based on the policies defined by the user. Midway through an agent session, the user might discover that the task requires a tool currently denied by the policies, or requires a legitimate but forbidden sequence of tool calls. A user might also realize that their policies allow an action that must be disallowed, or must be allowed under stricter conditions. In all such cases, the user must be able to update their policies mid-session. The engine must apply such updates, and also keep receiving events and judging requests, without any pauses: tool calls in flight must complete and be recorded, and if several agents share the engine, the other agents must continue undisturbed.
To this end, the engine supports dynamic policy updates: users can update policies without stopping the engine or the agent. Handling policy updates correctly requires care in the implementation. The first issue is that a new or modified policy may refer to events and information that the engine has not tracked. For example, as described in the previous section, the engine does not maintain a full record of all prior events: it prunes events from the log when their information has been incorporated into the state snapshot, which only tracks a subset of information needed to evaluate the current policies. Second, if requests arrive concurrently with policy updates, it is important that the request is evaluated against a consistent view of the policy set. Evaluating the request against a mixture of policies from before and after the update could lead to a verdict that incorrectly allows an action, which would be insecure. This atomicity of policy updates is also relevant when it comes to crashes: if a crash happens while the policy set is being updated, then after the system recovers, the policy set must be entirely the pre-update policy set, or it must reflect all of the effects of the update.
To deal with the first issue of missing information, the local engine adopts the convention that when a policy is updated, the history of events used for evaluating its temporal clauses only includes events that arrive after the policy is updated. The history for a new temporal clause does not retroactively include events the engine saw prior to the policy update. This avoids any problems related to the engine only having stored partial information about those past events. Going forward after the policy is updated, the engine will track the relevant information about subsequent events.
For the second issue about atomicity in the presence of crashes and concurrency, the local engine linearizes and persists policy updates in the same log as it uses for events. The position of a policy update in the log determines its relative order with respect to events: requests that follow the policy update record in the log will see the fully updated policy set. Similarly, when a crash occurs, the engine will replay the policy update records in the log in order to return to a consistent state.
Example
Suppose that a session began with a policy that permits git:commit unconditionally, and partway through the user decides that the agent must lint the code before it commits. The user can update that policy as follows (assume the schema declares git:commit, and a lint action whose response contains output.clean):
@id("commit")
permit (principal, action == Sandbox::Action::"git:commit", resource)
when temporal {
!Sandbox::Action::"lint"::response{ output.clean: false }
since within 15m
Sandbox::Action::"lint"::response{ output.clean: true }
};
The figure shows an event log where the engine linearizes and persists a policy update in the same fashion as events. At @300, the agent’s lint run finishes successfully with clean set to true. At @310 the engine records the policy update. At @320, the agent requests permission for a commit action, and the engine denies this request, even though a successful lint event occurred earlier. This is because the engine disregards prior events for the updated policy. For evaluating an updated policy, the engine only considers events that occur after the policy is updated. Thus, the agent must run lint again at @330, and once that run finishes successfully at @340, the engine allows the commit at @350.
The cost of a verdict
Because a tool call cannot start until the engine returns its decision, the decision time adds latency to every tool call. Part of the decision time is the sync: the engine writes every event to disk to ensure durability. The sync time depends on the storage device. The rest of the decision time is the evaluation of the temporal conditions that apply to the request. We discuss here two key schema and policy design decisions that influence the evaluation time: action granularity and window size.
Action granularity
Every policy in Dogwood has a scope, which specifies the principal, action, and resource the policy applies to. The scope is written in the permit line of the policy. Consider the scope of policy push_after_green_tests from the earlier example:
@id("push_after_green_tests")
permit (principal, action == Sandbox::Action::"git:push", resource)
when temporal { ... };
The policy’s scope is limited to the action git:push, which means that the policy has no influence on the verdict of any other action’s request. Thus, for a given request, the engine need only evaluate the policies whose scope covers the requested action: it evaluates push_after_green_tests on a push request and skips it on all other requests. How tightly a policy can be scoped depends on the actions we declare. When declaring the set of actions, our goal is a granularity at which each policy can be scoped tightly.
Suppose we had written the same policy under a coarser schema consisting of a single git action with subcommand as an input string:
type GitInput = { subcommand: String, args: Set<String> };
action "git" appliesTo {
principal: [Agent],
resource: [Box],
context: { input: GitInput }
};
@id("coarse_git_push")
permit (principal, action == Sandbox::Action::"git", resource)
when { context.input.subcommand == "push" }
when temporal { ... };
The policy coarse_git_push permits exactly the requests push_after_green_tests permits: the policy’s when clause rules out all subcommands other than push. Its scope, however, covers every git request, so the engine evaluates it on a status or a commit as well as on a push. Such evaluation is wasteful: the policy will never permit a status or a commit. Yet the engine only learns this after evaluating the policy. The original policy push_after_green_tests avoids this wasteful evaluation because its scope is limited to action git:push, and the engine evaluates it only on push requests.
To illustrate the performance impact of action granularity, we wrote the same policies under two schemas for git actions, one that is coarse-grained and one that is fine-grained. The coarse-grained schema declares one git action with the subcommand in its input. Using the coarse schema we wrote a hundred policies, twenty per subcommand for push, commit, fetch, checkout, and status, each selecting its subcommand in a when clause just like policy coarse_git_push. The fine-grained schema declares five actions, git:push through git:status. Using it we wrote the same hundred policies (twenty per action), each scoped to its action just like policy push_after_green_tests.
We then measured the engine’s evaluation time for push requests by simulating agent sessions of fifteen minutes to twelve hours. In every session, the engine reached the same verdicts under the two schemas, and evaluation under the fine-grained schema was roughly five times faster. Under the coarse-grained schema, a push request is in the scope of all hundred policies, so the engine evaluates all of them. Under the fine-grained schema, the push action is only in the scope of twenty policies and the engine evaluates only those, making this setting five times faster.
Window size
The window of a temporal condition is the time duration of the most recent history that the condition can refer to. Whether a temporal condition holds depends only on the events falling inside the window. The window can be specified in the within clause of the condition. For example, the condition of push_after_green_tests below sets a fifteen-minute window and the condition depends only on the events of the last fifteen minutes:
!tests::response{ output.passed: false } since within 15m tests::response{ output.passed: true }
We define window size as the number of events inside the window. To evaluate a condition and update the in-memory state, the engine may need to scan every relevant event in the window, so its evaluation time grows with the window size. To see how window size affects evaluation time, we wrote the condition above with a fifteen-minute window and with a 24-hour window and simulated an agent session producing 360 events per hour. We measured the engine’s evaluation time for a push request at session lengths from five minutes to twelve hours. The plot below shows the result. Each point in the plot is the median of 200 evaluations on one server-class host; repeated runs were within five percent of each other.
The solid line shows the evaluation time of the condition with the fifteen-minute window. Once the session grows past fifteen minutes, the window holds the latest ninety events no matter how long the session runs. Thus, the window size remains constant after fifteen minutes and the evaluation time plateaus. The dashed line shows the evaluation time of the condition with the 24-hour window. This window contains all events of the growing session, so the engine scans more events with every hour and the evaluation time keeps growing. At twelve hours, the engine takes six milliseconds per decision, three hundred times longer than under the fifteen-minute window.
In general, therefore, shorter windows can improve policy evaluation performance. On the other hand, sometimes a window cannot be made shorter because its length is important for the underlying requirements. For example, a requirement that limits refunds to a hundred dollars per day needs a 24-hour window, and no shorter window can express this constraint. Because changing the window changes the semantics of the policy, we must take care when shortening a window to save evaluation time: the shorter window must still capture the requirement.
Conclusion
Agents accomplish tasks through external tool calls. They run tools that change files, call services that place orders, and transfer data between systems. Left unchecked, however, these tool calls can have irreparable consequences. As agents scale to settings where they work autonomously for longer intervals with more tools, we need safeguards that can regulate how those tools are used. Dogwood allows users to write policies that restrict tool use, and the Dogwood Local Engine makes it easier for developers to integrate Dogwood policy enforcement into their agentic software.
The Dogwood Local Engine is available now on GitHub. In this post, we described just a few of Dogwood’s features; the language also supports other operators and primitives such as previous, count, and sum. We’re actively adding more constructs and developing both the language and the engine. We welcome feedback on both the language design and the implementation of the local engine.