On September 14, 2026, Perplexity added its local AI agent, Portable Computer, to the Windows version of the Perplexity app. Beyond the existing NVIDIA DGX Spark and RTX PCs running Linux, it now extends to compatible Windows PCs with GeForce RTX GPUs and to RTX PRO workstations.
But "Windows PC" here does not mean the widely available Copilot+ PCs or typical gaming PCs. On-device inference requires at least 24GB of VRAM, and users must also hold a Pro or Max plan for individuals or businesses. Windows support does broaden the base, but for now it is an expansion aimed at developers, creators, and enterprise workstation users who already own top-tier GPUs.
From RTX PCs running Linux to Windows
Portable Computer is a local-first version of Perplexity Computer, which plans and carries out multi-step tasks on the user's own machine. According to Perplexity, the Windows version lets users choose a local model from the model selector in the existing app, then download whatever model they need. As a concrete example, NVIDIA cites Qwen 3.8 27B, which has been post-trained for Perplexity Computer and optimized for RTX GPUs.
According to NVIDIA, the Windows release follows the already supported DGX Spark and RTX PCs running Linux. Because it works from the existing Perplexity app on Windows machines that meet the requirements, users can get started without setting up a Linux environment. Still, having more OS options is not the same as running on an ordinary PC.
It's not just the model: the runtime stays local too
It would miss the point to see this product as just another "local LLM app." According to Perplexity, more than the model runs on the device. A planner breaks work into steps, and a tool router chooses which capabilities to use. A deterministic orchestrator and an agent harness (the execution layer for agents) run approved tool calls in an isolated sandbox. A scheduler and persistent task queue keep recurring jobs and state intact, and a local search index looks up information on the device.
This combination decouples long-running work from the "ask once, get an answer" pattern. Recurring jobs, such as reconciling invoices in an approved folder every morning or triaging pull requests against a local codebase, need more than a model's reasoning ability. They require a queue that preserves state, an index, file permissions, and a tool execution environment. Tasks completed locally do not consume Computer credits, so work that repeats often can avoid metered cloud charges. In exchange, users bear the cost of the GPU, electricity, and maintenance.
The 24GB line: supported GeForce cards skew toward the top end
Of the five models checked against NVIDIA's standard specifications, the GeForce cards meeting the 24GB requirement are the RTX 3090, 3090 Ti, 4090, and 5090. The RTX 5080 (16GB) falls short.
| GPU | NVIDIA standard VRAM | 24GB requirement |
|---|---|---|
| GeForce RTX 5090 | 32GB | Met |
| GeForce RTX 5080 | 16GB | Not met |
| GeForce RTX 4090 | 24GB | Met |
| GeForce RTX 3090 Ti | 24GB | Met |
| GeForce RTX 3090 | 24GB | Met |
This table compares NVIDIA's standard specifications as of September 16, 2026, against Portable Computer's 24GB requirement. The figures were checked against the RTX 5090, RTX 5080, and RTX 3090 Family pages, and for the RTX 4090, Appendix A of the NVIDIA Ada GPU Architecture.
What the comparison shows is that memory capacity, more than how new the product generation is, acts as the hard cutoff. The RTX 5080, a high-end card of the current generation, is excluded, while the RTX 3090 from two generations ago qualifies. Tensor Core generation or advertised AI performance alone cannot tell you whether a card is supported. Note that this classifies only the five models whose official specifications were checked; it is not an exhaustive list of all GeForce products, RTX PRO cards, or custom OEM configurations. Meeting the VRAM requirement also does not guarantee that other conditions, such as the OS, drivers, and storage, are satisfied.
"Local" does not mean fully offline
Portable Computer defaults to local processing, but it is not a product that cuts off the outside world. It can bring Perplexity Search, connected services such as Gmail, Outlook, Slack, and GitHub, and multiple frontier models in the cloud into the same task. It also offers a setup that operates desktop apps through a local MCP server.
Perplexity and NVIDIA say that when information needs to leave the device, the user is asked for permission before processing. The privacy value, then, lies not in an absolute guarantee that "nothing leaves the machine," but in users being able to choose the boundary between local file processing and search, connectors, and cloud advice. For enterprise deployment, the next questions are not just the permission prompt but whether administrators can audit which pieces of data were sent to which service, and which retention rules apply at the destination. This announcement does not describe that level of granular control.
The 85.4% score on DGX Spark is not a Windows benchmark
To show the effectiveness of local execution, Perplexity published an in-house Local Knowledge Work Bench that simulates everyday knowledge work. Running 53 tasks three times each on DGX Spark, the combination of Computer and Qwen 3.8 27B scored 82.6%, the general-purpose harness Pi scored 77.6%, and Hermes scored 74.0%. PPLX 27B, a version of Qwen with additional training, reached 85.4%, according to the company.
On speed, however, Pi averaged 176 seconds, Computer with Qwen 218 seconds, and PPLX 27B 250 seconds, so the highest-scoring configuration was not the fastest. Token consumption was 520,000 for Computer with Qwen, versus 678,000 for PPLX 27B. Perplexity also acknowledges that even though Qwen 3.8 27B has a nominal context of 260,000 tokens, it empirically struggles beyond 100,000. For that reason, the design keeps the base prompt and basic tools small and loads only the needed capabilities as skills.
These results do suggest that the joint design of model and harness, not the model alone, shapes performance. But this is a small in-house benchmark designed and measured by Perplexity itself, not an independent reproduction. It was also run on DGX Spark, whose memory configuration and execution environment differ from a Windows PC with a 24GB GeForce card. Accuracy, runtime, power consumption, and actual VRAM usage for the Windows version have not been published, so the 85.4% cannot be read as a performance guarantee for Windows.
What recurring work needs: Windows measurements and control over outbound data
To hand off tasks like invoice reconciliation or codebase triage every day, runtime, power consumption, and actual VRAM usage on real Windows hardware will determine deployment costs. Moreover, whenever search or connectors are used, if users and administrators cannot trace what information left the device, it will be hard to verify in practice the benefit of choosing local processing. Neither point is adequately addressed in this announcement.
Support for GPUs with less than 24GB has not been announced either, and whether quantization or smaller models could lower the requirement remains undetermined. What will turn Windows support into broad adoption is not confirmation that it runs on top-tier GPUs, but publishing measurements for recurring workloads and controls over outbound data, and how far the range of supported machines can be widened.
