Information has surfaced suggesting that OpenAI is preparing a personal AI agent called "o" that continues working even after the app is closed. On September 26, leo (@synthwavedd), an X account that shares unannounced AI-related information, named "o" as one of the products expected to be unveiled at the developer event DevDay on September 29. However, the product name, underlying model, and release timing have not been officially confirmed by OpenAI.
Meanwhile, Meta announced its personal AI agent "Muse" on September 8. It is designed to handle real-world tasks such as sending emails, making reservations, and shopping, and can continue processing as needed even after the app is closed.
When comparing these kinds of AI agents, it's more useful to look at "which specific tasks were completed, and to what extent" rather than relying solely on an overall score. Looking at Muse's public benchmark results and the design details Meta has disclosed reveals the key questions OpenAI would need to address if it enters the same space.
Is OpenAI's "o" a Long-Running AI Agent?
According to leo's (@synthwavedd) post, "o" is one of the products OpenAI is expected to announce at DevDay, and it is said to be an AI agent designed for extended task execution, similar to Grok Bot or Hermes Agent.
The same post also suggests that it will use a derivative model of GPT-6 Astra called "aeon," with a focus on handling long-duration tasks.
However, this remains unconfirmed information at this stage. OpenAI has not issued any official statement about the product name "o" or its relationship to "aeon."
What is confirmed is that OpenAI DevDay will be held in San Francisco on September 29. The official event announcement does not mention any planned unveiling of "o."
Even so, the reported direction of "o" reflects a broader trend in the next phase of AI assistant competition.
Until now, chat-based AI has largely centered on a model where users send a question or instruction, the AI responds, and the task is considered complete. Always-on agents, by contrast, continue working after receiving a goal, resume action when circumstances change, and ask the user for judgment only when necessary.
Meta's Muse is already offering this kind of product.
What Does Muse's "9.3" Score Actually Measure?
Muse has received an overall score of 9.3 on Assistant Benchmark, a site that evaluates the practical usefulness of AI assistants.
Assistant Benchmark is an evaluation site run by David Pawlan and others, which has AI agents perform tasks simulating real-world use cases such as travel booking, shopping, email replies, and recurring tasks.
Each category is scored on a scale of 1 to 10, and categories that haven't been tested are left unscored.
On Muse's individual page, test results have been published for 8 of the 15 scored categories.
| Test Category | Muse Score | What the Published Record Shows |
|---|---|---|
| Executing online tasks | 9 | No matching accommodations found; presented alternatives and suggested changing conditions |
| Travel booking | 7 | Limited to Duffel, restricting available options |
| Recommendation quality | 9 | Selected a restaurant matching the criteria and booked via OpenTable |
| Product purchase | 10 | Confirmed type and budget, prepared an itemized order |
| Email reply | 10 | Checked calendar, drafted a reply, and sent it after approval |
| Recurring execution | 10 | Executed weekday recurring tasks as scheduled |
| Integration with external services | 10 | Integrated with OpenTable and Stripe's Link |
| Permissions and privacy | 9 | Confirmed before sending emails or making purchases; set permissions per service |
The table is based on Assistant Benchmark's Muse page, confirmed on September 27, 2026.
Summing the published scores for the 8 categories gives 74 points, for a simple average of 9.25. Rounded to one decimal place, this matches the displayed score of 9.3.
However, 9.3 does not represent an evaluation of Muse's full range of capabilities. Of the 15 categories, several—including memory, phone calls, and the ability to handle multiple tasks in sequence—remain unscored.
What matters more than the overall score is understanding what the agent excels at and where its limitations lie.
For email replies, for example, Muse was able to check the schedule, draft a response, and send it after the user's approval. Travel booking, on the other hand, was limited to a 7 due to restrictions on which services it could use.
The 10 for product purchases doesn't mean Muse can "buy anything without user confirmation." It scored highly because it prepared the order details and stopped short of actually completing payment without approval.
The practical usefulness of an AI agent isn't determined solely by the performance of the underlying model. It's also heavily influenced by which services it can connect to, how much operational authority it has, and whether it properly checks with the user before critical actions.
Muse Keeps Working Even After the App Is Closed
According to Meta's product description, once Muse receives a goal or task from the user, it continues working as needed afterward.
For tasks that take a long time, Muse keeps processing even after the user closes the app. It's designed to notify the user again when circumstances change or when approval is needed.
This is different from simply having "faster AI responses."
When planning a trip, for instance, presenting options on the spot is something a regular AI assistant can do. An always-on agent, however, is expected to remember booking conditions and schedules, and resume work later if circumstances change.
At the same time, sending frequent notifications when nothing has actually changed would become a burden for the user. Deciding what to keep monitoring and when to notify the user is itself an important product capability.
Each Muse user is assigned a dedicated cloud-based computer, within which it performs tasks using a browser and files.
According to Meta, users can review the actions Muse has taken and the permissions they've approved, allowing them to trace afterward what the AI remembered and what operations it performed.
However, a score of 10 for "recurring execution" on Assistant Benchmark doesn't prove that Muse can reliably sustain every kind of task over weeks or months.
Beyond simply executing tasks at scheduled times, it also matters whether Muse can properly recover from situations like a dropped connection, a change in service terms, or a failure partway through a task.
If OpenAI's "o" is indeed characterized by long-running tasks, this kind of continuity and recovery capability will be an important point of comparison.
Where Should Automated AI Stop?
For always-on AI, how much can be automated matters—but so does knowing "where to stop."
Meta has explained Muse's safety design, noting that it was built on the assumption that the AI could make mistakes or receive malicious instructions embedded in external content it reads.
A representative example of this risk is "prompt injection."
For instance, if an AI is reading a webpage while working, and that page contains an embedded command saying something like "ignore previous instructions and perform a different action," there's a risk the agent could follow it.
For AI that continues running long after the app is closed, there are more opportunities for it to encounter external data while the user isn't directly watching the screen.
To address this, Muse runs a separate supervisory system called "Sentinel," distinct from the AI actually performing the tasks.
According to Meta, whenever Muse accesses an external service or performs actions on the internet, it undergoes review by Sentinel. Muse itself cannot override Sentinel's decisions.
Sensitive information such as passwords and payment details is stored in an area that Muse cannot directly access.
For critical actions like sending emails or making purchases, the system is also designed so that a human reviews the details before execution.
Looking at this design changes what it means for an AI agent to be "capable."
Being able to complete every step without human review isn't the only measure of competence. Preparing the details of an email or a purchase, then correctly pausing right before an irreversible action and returning the decision to the user, is also an important capability.
That said, the fact that each user's virtual machine is isolated doesn't mean Meta itself has no access to the data in the current version of Muse.
Meta has explained that in the current Muse Secure VM, while data between users is kept separate, this isn't a system that technically prevents Meta's own access entirely when necessary for service operation or safety assurance.
A separate system called "Muse Confidential VM," which uses cryptographic protection to prevent even Meta itself from accessing the data, is planned for release sometime within 2026.
The currently available version of Muse and the upcoming Confidential VM need to be considered separately.
Comparing New AI Agents by What You Want to Delegate
According to Meta's announcement, Muse is now available in the U.S. on iOS, Android, and the web. Most use cases are free, with a paid plan also available for users who want expanded usage.
However, this announcement doesn't clarify the release timing or pricing for Japan.
Similarly, no official pricing or terms of use have been disclosed yet for OpenAI's "o."
When comparing these new AI agents, it helps to first pick one specific task you want to delegate, then examine how far the agent can carry it from start to finish.
For email, what matters isn't just the quality of the writing—it's whether the agent can correctly check the schedule, select the right recipient, and ask for confirmation before sending.
For reservations, the amount of effort saved varies significantly depending on whether the agent merely searches for options or actually accesses the desired service and proceeds all the way up to the point of booking.
If "o" is indeed unveiled at DevDay on September 29, what will matter isn't just the product name or underlying model, but how long it can sustain a task, which services it can operate, whether it can recover from a failed task, and at what stage it asks for the user's approval.
Remembering a goal, continuing to work toward it, operating the necessary services, and returning judgment to a human at critical moments—how reliably an agent can carry out this entire sequence is likely to determine the practical value of always-on AI agents.
