On September 8, NVIDIA outlined its development plans for "CUDA Rust," an effort to let developers write GPU kernels in Rust, and presented two approaches: "cuda-oxide," which works directly with threads, and "cutile-rs," which works with blocks of data. The work aims to take Rust beyond calling the GPU from Rust applications to writing the code that actually runs on the GPU. Both projects are in early development and are not ready for production use. Even so, their designs use types to prevent errors such as overlapping write destinations and mistakes tied to asynchronous execution, which could reduce the amount of checking GPU developers have to do by hand.

AD

From calling the GPU from Rust to writing GPU code in Rust

cuda-oxide extends the code-generation stage of the Rust compiler and converts GPU-targeted functions into PTX, an intermediate code for NVIDIA GPUs. The pipeline runs from Rust's intermediate representation through Pliron and LLVM, and lets developers write CPU-side code and GPU kernels in the same source file. The official repository also includes an example in which the host passes a closure to a generic kernel.

Its programming model is SIMT, widely used in CUDA C++. Developers specify which element each thread reads and where it writes. This allows fine-grained control over the execution configuration, but it requires an understanding of how threads and memory work.

cutile-rs, also known as cuTile Rust, writes work in larger units. A tile is a small region cut out of an array or matrix. Developers describe operations on tiles and leave the mapping to physical GPU threads to the compiler. A dedicated macro embeds the kernel's syntax tree in the CPU-side executable, and it is compiled through CUDA Tile IR on first run.

In the cutile-rs addition example, an output of 1,024 elements is split into chunks of 128. That split also determines the execution configuration: 1,024 ÷ 128 = 8 tiles. Instead of computing write positions from thread indices, the developer hands each pre-divided region to a unit of work. It is, however, a DSL for describing kernels, and arbitrary Rust programs do not become GPU code as they are.

There are precedents for writing GPU kernels in Rust. NVIDIA's ecosystem overview also introduces rust-cuda, which goes through NVVM, and CubeCL, which has backends for multiple GPUs. CUDA Rust can be seen as NVIDIA building a CUDA-oriented option on top of that body of work. Being able to choose Rust is not the same as being able to move the same kernel to another vendor's GPU.

Extending ownership until the GPU finishes executing

In Rust, a value can be read from many places at once, but a mutable reference and a read-only reference cannot coexist. Bringing these ownership and borrowing rules to the GPU raises a problem when many threads write to the same array at the same time.

cuda-oxide's DisjointSlice is a type for giving each thread a non-overlapping write destination. Instead of sharing a mutable reference to the whole array across all threads, each thread uses a dedicated index type to get access to the element it is responsible for. If the index is out of range, no value is returned, so the caller can handle that case.

Types alone do not make launch conditions automatically correct, however. The launch configuration is validated in advance against the conditions declared in launch_contract, and the verified values are passed to a safe launch API. Paths that use a raw launch configuration require unsafe. The design combines compile-time borrow checking with pre-launch configuration validation. (cuda-oxide launch API description)

cuTile Rust splits the tensor to be written on the host side and gives each tile an exclusive region. A tensor is a multidimensional array, and the distinction between shared inputs and mutable outputs carries into the kernel. Developers do not manipulate physical threads directly, which is one of the conditions that makes the safety guarantee possible.

Also, the GPU keeps working after the CPU launches a kernel. If the CPU-side function returns and the array can be freely modified at that point, it could collide with GPU reads and writes. The cuTile Rust research paper "Fearless Concurrency on the GPU" brings this timing gap under ownership management as well, through DeviceOp, which defers execution.

A sequence of operations can be connected to synchronous execution, Rust async execution, or re-execution as a CUDA Graph. In the safe API, ownership is returned to the caller only after the GPU work completes. What is protected is the write destinations inside the kernel and the lifetime of data while the GPU is running.

The guarantees have limits. In cuda-oxide at the time of the announcement, using shared memory requires unsafe, and cuTile Rust also retains paths that step outside the safe tensor API for low-level operations. Using Rust does not mean every kind of GPU code error is eliminated.

AD

Before you try it: differences between the announcement and the current README

For cuda-oxide, the announcement blog gives the setup requirement as CUDA 12.x or later, while the README as checked on September 13 says CUDA Toolkit 13.0 or later. Comparing the SIMT section of NVIDIA's September 8 blog post with the Setup section of the cuda-oxide README also shows a difference in the pinned Rust nightly date.

Item cuda-oxide cutile-rs
Unit of work described Thread-level SIMT Tiles, small regions of data
Rust requirement Current README specifies nightly-2026-08-28 Stable Rust 1.89 or later
CUDA requirement Current README: Toolkit 13.0 or later, R580 or later driver supporting CUDA 13.x Current README recommends CUDA 13.3
OS Linux Linux
Setup caveats Announcement blog states CUDA 12.x or later and nightly-2026-04-03 Supported GPUs vary by CUDA version
Development status Early alpha Early development; not ready for production use

The table reflects public documentation for both projects as checked on September 13, 2026. It does not establish when or why the version differences arose, nor is it the result of testing compatibility on real hardware. When setting up, the version of the code you fetch needs to match the environment that version specifies.

The supported-GPU description for cutile-rs can also be misread if taken simply as "a newer CUDA is required." sm_100 and later are supported from CUDA 13.1, sm_8x was added in 13.2, and 13.3 adds sm_90, covering sm_80 and later. These are architecture numbers that indicate a GPU's compute capability.

cutile-rs imposes a lighter burden in that you can start with stable Rust. You still need to check separately whether your GPU has a compatible CUDA environment. cuda-oxide, for its part, requires clang and the libclang development headers in addition to the Rust toolchain. Ease of writing programs and the effort of assembling the environment should be evaluated separately.

What the performance research shows, and what can't be entrusted yet

How much GPU performance does cuTile Rust's safe style of programming preserve? The preprint "Fearless Concurrency on the GPU" by Melih Elibol and colleagues evaluates the design on matrix multiplication on a B200. It is not a study that measures the speed of CUDA Rust as a whole.

In the paper's safety-overhead evaluation, half-precision matrix multiplication was measured on a single GPU in a DGX B200. With each matrix dimension at 8,192, cuTile Rust reached 2.07 PFLOPS, or 96.4% of cuBLAS. The difference from an unsafe version using raw pointers with the same processing order was reportedly within 0.3%.

The measurements fixed the SM clock and tuned tiles for each matrix size. The initial JIT compilation cost is also treated separately from steady-state performance. What the authors' results support is the assessment that the mechanism for ensuring safety added little overhead under these conditions, not a guarantee that Rust achieves the same performance on every workload.

A more practical example is Hugging Face's Grout, an experimental engine for trying out inference for Qwen3, which uses kernels written in cuTile Rust. However, it uses cuBLAS for the matrix computation in linear projections. It uses cudarc for calls from Rust, so it is not a configuration in which all GPU computation has been replaced with custom Rust kernels.

The paper also explains that Grout uses unchecked access and raw pointers in attention and in normalization that fuses several operations. Simple operations use the safe API, while developers take on responsibility for processing that cannot yet be expressed safely. There is a gap between being able to build this into a real inference engine and being able to write every path in safe Rust alone.

NVIDIA plans to keep maturing CUDA Rust into 2027 and beyond and has not given a date for a stable release. For developers considering adoption, the deciding factors are whether the operations they need can be expressed in the safe API, whether the remaining unsafe portions can be verified, and whether they can pin the GPU and toolchain they use and try it out. If those conditions are met, they can test for themselves, on their own workloads, the benefit of using Rust to tie together data management across the CPU and GPU sides.