If 100 million people used Meta's personal AI agent Muse, it would need about 1.58 million CPUs and 800PB of memory. That is the scale of estimate Wccftech published on September 21, 2026. Because Muse gives each user a dedicated virtual machine and keeps working after the app is closed, it offers a concrete example of how the spread of AI could change demand for CPUs and memory.
But this figure is not a Meta procurement plan. It is a simple sum of per-user resource allocations. To work out how much hardware is actually needed, you have to separate how many AI agents run at the same time from how much data must be kept while they sit idle.
What do the 800PB and 10,000PB figures mean?
Wccftech's estimate assumes each user is allocated 2 vCPUs, 8GB of RAM and 100GB of SSD storage. A vCPU is a unit for the CPU resources available to a virtual machine. The article cites a post on X, but I could not find an official document confirming that Meta guarantees this configuration to every user. The 100 million user figure should likewise be read as an assumption in the estimate, not an official result or target.
Under these assumptions, the simple math looks like this. Capacities use decimal units: 1PB is 1 million GB, and 1EB is 1,000PB.
| Item | Assumption per user | Calculation for 100M users | Total |
|---|---|---|---|
| Virtual CPU | 2 vCPUs | 100M × 2 | 200M vCPUs |
| RAM | 8GB | 100M × 8GB | 800PB |
| Storage | 100GB | 100M × 100GB | 10,000PB (10EB) |
Allocating 8GB to each of 100 million people gives 800PB, and 100GB each gives 10,000PB. But these are totals of what is allocated to users. They do not mean Meta must buy 800PB of RAM or 10EB of SSD outright.
The original article also contains internal inconsistencies. As of September 22, its body text gives the RAM figure as 100PB, which conflicts with the 800PB in the headline. 8GB × 100 million users corresponds to 800PB. In addition, the article gives about 1.58 million CPUs and then says that if only 10% were active it would be 15.8 million. Proportional scaling under the same conditions would give about 158,000, so the order of magnitude is off.
Even after correcting these errors, there is a large gap between resources allocated to users and resources actually in use at the same time. A user allocated 100GB will not necessarily use all of it. A virtual machine with 2 vCPUs is not necessarily running its CPUs flat out at all times. Accounting for these differences is the starting point for turning a huge allocation total into real hardware demand.
A dedicated VM is not the same as a dedicated physical CPU
AWS technical documentation distinguishes between a vCPU that maps to a physical core and one that maps to a hardware thread under SMT. SMT is a mechanism that lets one physical core handle multiple threads of execution. So to convert 200 million vCPUs into a number of physical CPUs, you first have to decide what counts as one vCPU.
Wccftech's figure of about 1.58 million CPUs can be roughly reproduced by dividing 200 million vCPUs by 126 cores. However, the basis for the "126-core" CPU the article cites is unclear. The AMD EPYC 9754, as published by AMD, has 128 cores and 256 threads. I can't tell whether 126 reflects cores reserved for management or is simply a typo. The name "Ryzen EPYC" also conflates two different product lines.
Assuming a 128-core CPU and one vCPU occupying one physical core, 200 million ÷ 128 gives 1.5625 million CPUs. If one vCPU instead maps to one thread and you divide by 256, you get 781,250. But the latter only changes the unit being counted; it does not mean the same work can be done with half the CPUs. Actual processing performance depends on what the AI agents run.
Moreover, even a running virtual machine is not always using the CPU. Some of the time it is waiting for a response from an external site or for the result of AI model inference.
To estimate CPU demand, it is easier to see the required variables by multiplying the number of users by the share running simultaneously and by the average physical-core equivalent consumed per active VM, then dividing by the processing capacity one CPU can provide.
As an illustrative formula:
Number of CPUs = users × concurrent-activity share × average cores consumed per active VM ÷ (cores per CPU × target utilization)
A design that pushes every CPU core close to 100% at all times tends to create processing queues when traffic spikes, so headroom is needed to meet the desired response time. This is not a formula for Muse's actual internal architecture; it shows variables missing from the original article's simple calculation.
Giving each user a dedicated VM and giving each user dedicated physical CPU cores are separate design decisions. Sharing CPUs among multiple VMs may reduce the number of physical CPUs required, but how far sharing can go while keeping performance adequate can only be known by measuring real workloads.
Estimate memory and storage separately
Even if concurrent CPU activity is 10%, RAM demand does not automatically fall from 800PB to 80PB. If an idle VM keeps its OS and in-progress state in memory, that capacity must still be reserved. Spreading out CPU usage over time is a different matter from freeing memory.
Suppose, for example, that 10% of 100 million users are active and use 8GB each, while the other 90% of idle VMs each keep 1GB resident. The RAM needed would be 80PB + 90PB, or 170PB.
The 1GB idle figure is an illustrative assumption, not a measured Muse value. Still, it shows why the concurrent-activity share alone cannot determine the memory required. How small idle VMs can be made becomes more important as the user base grows.
Mechanisms that move in-memory state to storage are already in practical use in the cloud. With Amazon EC2 Hibernate, RAM contents are saved to the EBS root volume and read back on resume so processing can continue. No instance usage charges accrue while hibernated; you pay for the storage capacity. It is a concrete example of separating the retention of state from keeping the CPU running.
However, Meta has not announced that Muse uses this approach. A hibernated VM cannot run tasks as it is, so a mechanism is needed to resume it at scheduled times or on external events. Frequent wake-ups increase the work of reading state back, while long hibernation lengthens response times. Designs that save memory have to be weighed against the wait times users will tolerate.
The 10EB of storage should also be distinguished from the amount of data users actually store. If average usage were 10GB, data for 100 million users would total 1,000PB, or 1EB. This too is an assumption to give a sense of scale, not a figure for Meta's real usage.
Real facilities also need space for the OS and management data, plus replicas and backups to guard against failures.
Meta's design explanation says data inside the VM is continuously backed up. Neither provisioning the available 100GB per user as physical SSDs nor sizing facilities based only on files users have saved is accurate. For memory, the question is how much must be held at the same time; for storage, it is what is saved, and in how many generations and copies.
CPUs are used for more than AI inference
Muse's VMs have jobs such as running a web browser and executing code. In the architecture Meta has published, alongside each user's work area there are services that handle credentials and a monitor called Sentinel that watches operations and communications, and persistent app state is stored in Postgres. AI model inference, meanwhile, is run by connecting to infrastructure outside the VM. The per-user VM does not handle all the computation the AI requires.
This division of labor also explains why AI's spread is drawing fresh attention to CPU demand. After an AI model decides what to do next, programs must actually be run, their results read, and the next action taken.
When you hand the AI a task like arranging a trip, for example, it does not just generate text. It opens booking sites, enters conditions, reads search results and saves the information it needs. If a task runs for a long time, the working state must be kept for that whole period.
So speeding up or cutting the cost of AI inference alone does not determine the total operating cost of a personal AI agent. Time running VMs and storage capacity for data are also needed, and if a task fails and has to be redone, consumption of both rises.
To evaluate the efficiency of the whole service, you need to look not at the amount of text generated but at how much computing resource and time it takes to complete one task a user has asked for.
It is also hard to tie CPU and GPU demand together at a fixed ratio. A request to draft a short document and a request to browse many websites for a long time differ in the ratio of computing resources spent on AI inference versus browser processing.
If Muse's user base grows, large volumes of these new kinds of processing could arise. But converting that into units sold of a particular CPU requires information about what tasks users actually hand to the AI.
Meta's CPU procurement and the conditions for adoption
Meta had been broadening its CPU sources before releasing Muse. In its April 24 announcement, AWS said deployment under its agreement with Meta would begin with tens of millions of Graviton cores. The scope covers a range of Meta workloads including AI, and multistep AI agent processing is listed among the uses.
What is given here is a number of cores, not CPUs, and it is not an order dedicated to Muse.
AWS's explanation also includes that Meta can run its own virtual machines on Nitro-based bare-metal environments. In other words, it secures physical server resources in bulk and builds its own service platform on top. Multiplying consumer cloud-PC pricing by 100 million users would therefore not give Meta's real operating cost.
Meta is also positioned as a lead partner and co-developer of the Arm AGI CPU that Arm announced on March 24. The idea is to optimize the infrastructure behind Meta's apps and combine it with Meta's own AI accelerator, MTIA.
All of these efforts were under way before Muse's September release, so the agreements cannot be read as resulting from Muse's growth after launch. Still, they confirm that Meta had already been investing in CPU infrastructure for running AI agents.
The September 8 Muse announcement also said Muse can be used from WhatsApp. That gives people an entry point for handing tasks to the AI from an app they already use, though the timing for availability in Japan is undecided.
As users increase, what matters for capacity planning is less the number of sign-ups than how many users' tasks run simultaneously during peak hours.
The key figures for gauging the infrastructure Muse needs are peak concurrent activity, the memory held while idle, and average storage use per user. Looking also at the time from request to completion shows whether saving computing resources has come at the cost of longer waits for users.
Whether Meta can build a system that keeps idle costs low while resuming work immediately when needed will be a key condition for deploying AI agents with dedicated per-user work environments at scale.
