On August 31, Broadcom announced VMware AI Factory as a software-defined platform for VMware Private AI Cloud. The company explicitly named a configuration using AMD Instinct GPUs and ROCm, bringing together GPU resource pooling, model sharing, and visibility into token throughput and latency under a single operational platform. AMD support itself is not new. What changed this time is that the choice of accelerator has been tied into a system for managing usage and costs after deployment.

Broadcom claims it can shrink the time from bare-metal server to first model deployment from weeks to hours. However, the company has not disclosed the reference configuration, comparison method, third-party verification, or actual customer measurements. Whether costs actually fall depends less on which GPU brand is used than on whether these conditions can be confirmed.

AD

A Broader Operational Scope Than Just AMD Support

AMD support did not appear suddenly this summer. In August 2025, Broadcom and AMD announced plans to develop a joint platform combining VMware Cloud Foundation (VCF), AMD Enterprise AI software, and AMD Instinct GPUs. Subsequently, VCF 9.1 made general availability support for AMD Instinct MI350 Series GPUs on May 12, 2026, listing PyTorch, vLLM, and OPEA as supported frameworks.

Breaking down the timeline: 2025 was the planning stage for the joint platform, May 2026 was the product-level support for running AMD GPUs on VCF, and August 31 was the announcement that bundles this support into an operational workflow spanning through to model deployment. Conflating the start of AMD adoption with the expanded management scope announced this time makes it harder to see where the actual cost-reduction mechanism lies.

August's VMware AI Factory announcement is about folding that GPU support into the operational pipeline. According to Broadcom's explanation, on certified AI ReadyNodes, the platform automates deployment all the way from bare-metal servers through vSphere, vSAN, and Kubernetes in one motion. Into this, the AMD GPU Operator and AMD DVX driver are integrated, connecting through to model serving. Users select their AI software and accelerator configuration, and VCF is said to deploy and manage inference and RAG workloads across virtual machines, containers, and GPU resources.

Supported servers include those from Cisco, Dell Technologies, Lenovo, and Supermicro, and VCF customers reportedly can run more than 150 open-source and commercial models. What has expanded here is not simply the capability to recognize AMD GPUs. Rather, it is that configuring systems, serving models, and managing post-deployment operations have all been placed under the role of a single product.

Where Do Sharing and Visibility Actually Affect Costs?

The entry point for cost management isn't how many GPUs you buy, but who uses already-allocated resources, and how much. VMware AI Factory pools and shares GPU resources, and also touts model sharing across tenants. If organizations can avoid deploying the same model separately for each department and holding excess GPUs as a result, redundant resource allocation could be reduced.

Model sharing allows the same model foundation to be accessed from isolated namespaces. This does not mean complete physical environment separation for each department, so organizations with strong data or regulatory requirements need to first determine the scope of what can be shared.

However, this benefit doesn't materialize automatically. The amount saved depends on GPU utilization rates, tenant isolation requirements, and contract terms. Adopting companies themselves must distinguish between redundancy that can be eliminated through sharing and the surplus that must be retained for isolation purposes.

Broadcom cited token throughput, latency, compute resources, and memory utilization as measurement metrics. Looking at these can provide clues for correlating processing volume with resource usage per model, helping identify over-allocation or congestion. Additionally, AI Gateway was presented as a feature that spans on-premises and cloud environments, handling routing and usage limits. Being able to measure usage and change routing as needed could serve as a tool for making costs more predictable.

That said, AI Gateway is listed among new or forthcoming features, and its general availability date was not disclosed in the announcement. Numbers appearing on a dashboard and actual cost reduction are not the same thing.

AD

The Real Difference from NVIDIA Lies in Contracts, Not GPU Brand

The previous VMware Private AI Foundation with NVIDIA was a separately sold add-on on VCF announced in 2024, requiring a separate license for NVIDIA AI Enterprise as well. Speaking to The Register, Broadcom's Prashanth Shenoy described VMware AI Factory as an evolution of this earlier product, supporting multiple accelerator configurations and AI toolchains. The official announcement also explicitly stated that users can choose their accelerator configuration, naming the AMD configuration specifically.

AMD executives stated in the same official announcement that the configuration combining AMD Instinct MI350 Series GPUs, ROCm, and VCF does not involve per-token usage-based billing. This condition differs from pricing models where costs rise with increased token processing. However, this does not necessarily mean the total cost of ownership—covering hardware, supporting software, power, and maintenance—is lower overall.

The fact that ROCm is an open software foundation also does not guarantee that CUDA-targeted code or NVIDIA-specific libraries can be ported without modification. Nor does the material provided include figures on what percentage of code can be migrated, or performance/cost comparisons of running the same model on both GPU brands. If migration work increases, any price advantage gained on the GPU side could be absorbed by increased software-related costs.

The central point of comparison is shifting away from the label of "AMD vs. NVIDIA" and toward which software and support are included in the contract, and where additional costs begin. The current announcement does not disclose individual pricing for VMware AI Factory, nor the scope included in VCF contracts. It also remains unclear whether operational features and licensing terms are equivalent between AMD and NVIDIA configurations.

Six Gaps That Leave Cost Verification Unresolved

Broadcom has presented features for managing costs, but has not disclosed individual pricing, the scope included in VCF contracts, release dates for each service, the comparison conditions behind the claimed deployment-time reduction, measured cost-reduction rates from actual customers, or whether operational conditions are equivalent between AMD and NVIDIA.

These six gaps are exactly why the published announcement doesn't allow for a numerical comparison of deployment benefits. First, regarding the claimed reduction from weeks to hours, we need to know what work is included in that timeframe and what reference configuration was used to measure it. Second, without clarity on when AI Gateway, the secure AI sandbox, or cross-tenant model sharing will actually become available, these features cannot be incorporated into operational planning.

What customers ultimately want to know—beyond what runs on a given GPU configuration—is how much costs will actually change once utilization rates, isolation requirements, and licensing are factored in. AMD support was already generally available as of VCF 9.1 in May 2026. The value of VMware AI Factory hinges on how far it can standardize deployment and operations beyond that point, and whether it can demonstrate actual measured cost differences.

Once pricing sheets, release dates, comparison conditions, and customer-measured figures are all available, companies will be able to judge how much automation and sharing can reduce GPU surplus in their own configurations. Until then, what Broadcom has presented is a set of tools for managing costs—not the savings themselves.