On September 30, 2026, Restate announced a $20 million Series A. Singular led the round, with participation from Redpoint Ventures and Capital One Ventures.
Restate was built by a team that worked on Apache Flink. Its goal is to make "durable execution" more widely usable. Durable execution records the progress of long-running AI agents and backend processes so they can resume partway through after a failure.
AI agents repeatedly query models and operate external tools. When a failure strikes during a process that runs for tens of minutes or several hours, cost and processing time depend heavily on whether everything must be redone from the start or whether the agent can pick up from completed work.
Looking at how Restate works shows where the boundary lies between what the platform can handle in failure recovery and what developers must still design themselves.
Why Preserving an Agent's Progress Matters
In June 2024, Restate announced a $7 million seed round and the release of Restate 1.0.
The problem the company originally tackled was running backend processes that span multiple services and APIs without losing intermediate state when failures occur.
This funding round can be seen as an effort to extend that mechanism to a broader range of backend processes, including AI agents.
Consider an agent that reads an order, checks inventory, waits for a manager's approval, and then processes payment.
If the connection to the server drops just before the final payment, restarting the whole process means re-fetching everything gathered so far, including AI model responses and inventory lookups.
To prevent this on the application side alone, developers would have to build their own mechanism to save completed work, check on retry how far the process had gotten, and resume from the right point.
Restate records model responses, tool execution results, and similar data in an execution history.
When a failure occurs, it restarts the process, but for steps that are already complete and stored in the history, it uses the saved results instead of performing the actual work again.
This reduces the time and cost of regenerating model responses that were already obtained or repeating work that was already finished.
However, this does not mean that every operation that failed before its result was written to the history is guaranteed not to run again. Official explanation for AI agents
Developers write processes as ordinary functions and wrap operations whose results may change from run to run, such as HTTP requests and model calls, in the SDK's ctx.run.
Restate records the result and reuses the saved result when recovering from a failure.
In other words, simply running an existing program on Restate does not automatically give every operation a recovery guarantee. Developers must specify which operations are to be recorded.
The official LangChain documentation likewise describes a model in which developers choose which parts of a tool call to make durable. How steps are recorded
Waiting for human approval works on the same principle.
The system issues an identifier for receiving the approval result and suspends the process while waiting for a reply. When approval arrives, the process is resumed using that identifier and continues from where it left off.
Because a function doesn't need to keep running while waiting, serverless environments avoid tying up compute resources just to wait hours or days for approval.
Of course, the absence of compute charges during idle time is a separate matter from Restate's storage and service fees, which are not free.
How Far Does the "Exactly Once" Guarantee Extend?
Operations on external services require particular care in durable execution.
Suppose a charge to a payment API succeeds, but the connection drops immediately afterward and the success response is never received.
At that point, the payment service may have already completed the charge, yet Restate's execution history may not contain the success result.
Retrying from that state could send the same payment request again.
Restate's official guide also says designs must anticipate the case where an external operation succeeded but its confirmation was never received.
Being able to reuse saved results is a different guarantee from ensuring the same operation is not executed twice on an external API.
The latter requires idempotency on the external API side.
Idempotency is the property that sending the same request any number of times with the same identifier does not duplicate charges, reservations, or similar operations.
Based on Restate's official specifications as of October 1, 2026, here is what the platform can manage and what remains with developers:
| Target | What Restate handles | What remains with the developer |
|---|---|---|
| Recorded model responses and results | Reuses saved results when recovering from a failure | Specifying which operations with variable results to record. Operations that fail before being recorded may be retried |
| Virtual Object state | Serializes writes to the same key and retains state | Different keys run in parallel. It is not a mechanism for exclusive control over an entire external database |
| Waiting for human approval | Holds the waiting state and resumes when a reply arrives | Designing the conditions under which approval is requested and who has approval authority |
| Charges and reservations on external APIs | Records call results and retries as needed after failures | Preventing duplicates on the API side with the same idempotency key, and implementing cancellation or compensation as needed |
The table is based on the specifications for recording steps, service types and state, approval waiting, and idempotency and compensation.
When using external APIs, it is important to use the same identifier on every retry and for the API provider to be able to reject duplicate requests based on that identifier.
Nor can every failure be solved by retrying.
An inventory shortage, for example, is different from a network failure. No amount of retrying will produce more stock.
If a different item has already been reserved or a fee has already been paid, "compensation" is needed to undo those actions.
Developers must write that compensation logic themselves to fit their business rules.
Restate provides a mechanism for executing processes without losing progress, but it does not decide business questions such as whether an order should be accepted or what should be undone on failure.
Combining State Management and Recovery Around the Execution History
In Restate's internal design, a replicated log serves as the foundation for recording and recovering processes.
According to a design explanation published in 2025, executed operations and their results are recorded in the log, and application state is managed based on that information.
State data is reflected in RocksDB, and periodic snapshots are stored in object storage.
Even if the Restate Server handling a process is lost, it can be recovered using the saved state and log.
The services that run the actual application logic are separate from the Restate Server, which manages execution history, state, retries, and more.
The Restate Server itself does not require a separate external database or distributed log system, bundling the necessary functions into a single system.
The aim of this design is to reduce the number of places where queues, state management, retries, and scheduling are spread across multiple infrastructure products.
Of course, the state that needs managing does not disappear. Restate takes over and consolidates part of the state that applications and multiple infrastructure products previously handled individually.
"Virtual Objects" can be used for conversation sessions and AI agent state management.
A Virtual Object holds state per key and executes only one write operation at a time for the same key.
This makes it easier to prevent problems such as multiple processes rewriting the state of a single conversation session simultaneously and producing inconsistencies.
On the other hand, operations on different keys can run in parallel, so one user's conversation can be processed while another's proceeds at the same time.
Long-Running Processes Also Need Code Versioning
For processes that run for hours or days, the application code may be updated while they are executing.
Here too, simply switching to the latest code is not always the right answer.
Restate ties a started process to a specific, unchanging deployment version. Even when retrying after a failure, it generally sends the process to the same version of the code.
This is because adding recorded steps mid-execution or reordering them could cause the stored history and the new code's sequence of operations to disagree.
Even after a new version is released, in-progress processes need to be able to finish on the old version that started them.
For long-running AI agents, consistency with code updates is an operational challenge alongside failure recovery.
Replit Uses More Than 10x the Durable Actions of the Previous Generation
In its funding announcement, Restate described a case in which Replit migrated the execution platform for "Replit Agent" to Restate.
According to Restate, the new architecture uses more than 10 times as many durable actions as the previous generation.
The "10x" here does not mean Replit Agent runs 10 times faster.
It means the number of units of work recorded inside the AI agent in a form that can be reused after a failure has grown substantially.
Rather than making the whole agent durable as one large process, the approach records fine-grained internal operations such as model queries, tool executions, guardrail checks, and state changes.
Finer recording units shrink the scope of work that must be redone after a failure.
If the 21st step fails after 20 have finished, the agent can resume near the 21st using saved results instead of redoing the first 20.
It also becomes easier to trace which operation had a problem.
On the other hand, more recorded operations means writing to the execution history each time.
Whether durable units can be made fine-grained therefore also depends on whether the latency and cost of recording are small enough.
What Restate emphasizes in the Replit case is a design in which durable actions are low-latency and low-cost enough to be built into an AI agent's inner loop.
According to Restate, Replit embeds thousands of durable steps in a single agent run, keeping the added latency per operation to a few milliseconds.
However, the "10x" figure in Restate's funding announcement comes without the absolute number of actions compared, the measurement period, or failure rates before and after migration.
This number alone therefore cannot be used to calculate how much processing speed improved, how much operating costs fell, or how much the agent's success rate rose.
For companies considering adoption, what matters is how finely they want to record their own agent processes and whether they can accept the added latency and cost.
Competing With Temporal: The Cost of Making Fine-Grained Steps Durable
Restate is not the only company seeing growing investment in durable execution.
In September, Temporal announced a $550 million Series E at a valuation of $12.55 billion.
Large amounts of capital are flowing into execution platforms that let AI agents and long-running backend processes continue through failures.
However, Restate is at Series A and Temporal at Series E, so the companies are at very different stages of growth. Funding amounts and valuations cannot be used to compare product performance.
What Restate stresses as a differentiator from existing durable execution platforms such as Temporal is latency and cost per unit of work.
If the burden of each durable operation is heavy, developers will record multiple operations together as larger units.
Conversely, if recording each operation is cheap enough, even fine-grained agent internals such as model calls and tool executions can be made durable.
Restate aims to make durable execution light enough to be built into ordinary backend processing, not just a few critical workflows.
One offering along these lines is BYOC (Bring Your Own Cloud), which lets customers use Restate in their own cloud environments.
BYOC Places the Runtime Inside the Customer's Cloud
On July 7, 2026, Restate announced the availability of BYOC.
With BYOC, Restate's runtime is deployed inside the customer's own AWS or Google Cloud account and VPC, and Restate manages it.
Compute, storage, and application data stay in the customer's cloud environment.
The aim is to make it easier to use Restate as a managed service even for companies that find it hard to move sensitive data, such as financial information, medical information, and source code, outside their own cloud.
Pricing is also based on the processing capacity reserved in advance, rather than pay-per-use billing on the number of actions processed.
In an estimate Restate presented in July, a workload averaging 5,000 actions per second, with headroom to handle 10,000 actions per second, would cost about $30,000 per month for the BYOC license and cloud infrastructure combined.
For the same average of 5,000 actions per second on Temporal, Restate estimated activity charges alone at about $300,000 per month, a difference of roughly 10x.
However, this comparison was prepared by Restate itself based on pricing terms as of July 2026.
For Temporal it calculates only activity charges, and it is not a third-party measurement of both products run under the same conditions.
Nor does it mean the difference will be 10x for every workload or contract.
Pricing based on reserved capacity tends to be more efficient for services that sustain high loads.
For services with long periods of low usage, cost-effectiveness depends on how much of the reserved capacity is actually used.
When comparing the total cost of an AI agent, one must also consider model inference, the compute that runs the application, and storage, not just Restate or Temporal fees.
Even if making fine-grained operations durable reduces retries, the evaluation needs to account for the total, including how much the recording itself costs.
AI Agents Need to Keep Going After Failure, Not Avoid It
Even as AI models themselves improve, it is difficult to make an agent process that runs for hours never fail.
Networks drop, external APIs time out, and servers stop. Sometimes a process waits hours for human approval.
For long-running AI agents, the priority is less about completely preventing failures than about being able to resume from where they left off without losing completed work.
Restate provides the mechanisms for this: execution history, state management, retries, and waiting and resuming.
On the other hand, idempotency to prevent duplicate processing on external APIs, compensation to undo partially successful business operations, and the decision of which operations to record still fall to developers.
To turn Restate's Series A into a real adoption decision, companies need to confirm, by inducing failures in their own agents, where processing can resume, how far duplicate requests to external APIs can be prevented, and how much latency and cost that adds.
If they can verify all that, developers can write less management code for failure recovery and spend more time on the core application design question of which tasks to entrust to AI agents.
