Documentation / EnglishView source ↗

Ruda — Rust High-Performance Computing

Rust Language Forks Issues Last commit

Website · Documentation

Full-stack API index · Tensor recipes · Backend composition · Data and storage · 中文文档

English | 简体中文 | 日本語 | Deutsch | Русский

Ruda is a Rust high-performance computing library, building a complete software stack from GPU kernels, compilers, and runtimes to mathematical computing, tensors, and models.

Ruda is building Rust compilation and execution paths targeting PTX, HIP, and custom ISAs, while retaining the CUDA C++ compilation path. Controlled low-level unsafe encapsulation, combined with Rust's type system, ownership, and borrowing at higher levels, balances low-level performance control with higher-level memory safety.

Quick Start

Requires Git, Rust/Cargo, a linker toolchain, an NVIDIA GPU and driver, and the CUDA Toolkit. See environment setup for installation details.

Use a published crate

cargo add ruda --features cuda

Add the CUDA backend to your application's Cargo.toml:

[dependencies]
ruda-driver-cuda = { version = "0.1", features = ["direct-ptx"] }

The examples below run from a source checkout.

Clone

git clone https://github.com/shuqi2077/RUDA.git
cd RUDA

Run a GPU kernel

Select the direct PTX compiler in your shell:

# Bash
export RUDA_CUDA_COMPILER=ptx
export RUDA_PTX_VERSION=8.0
# PowerShell
$env:RUDA_CUDA_COMPILER = 'ptx'
$env:RUDA_PTX_VERSION = '8.0'

Then build and run the example:

cargo run --release --locked -p ruda-driver-cuda --features direct-ptx --example ptx-runtime

The example runs FP32 addition on the GPU and prints PASS lines and compilation-cache counters. Select a PTX version supported by your GPU and driver.

Generate text with ruLLM

Place a local Qwen3.5-0.8B model in ./models/qwen35, or replace the path below with your model directory. Model files are not included; see model setup.

cargo run --release --locked -p ruda-llm --features nvidia-ptx --example qwen35_generate -- ./models/qwen35 "The capital of France is" 8 1

The example prints the generated text and token IDs. To use the CUDA C++ / NVRTC path instead, set RUDA_CUDA_COMPILER to nvrtc before running either example.

Use the native PyTorch backend

ruda-torch registers the PyTorch device ruda:0 on a single NVIDIA GPU. With PyTorch and setuptools installed and a C++20 compiler available, run from the repository root using the PTX environment settings above. On Windows, use an x64 MSVC developer shell.

cargo build --locked -p ruda-torch-native
python -m pip install --no-build-isolation --no-deps -e ./ruda-torch/python

The default loader finds this debug build automatically. Set RUDA_TORCH_LIBRARY to the library path when using a release build or another location. The Rust library and C++ extension must both use ABI 10; rebuild them together.

import torch
import ruda_torch

x = torch.arange(4, dtype=torch.float32).to("ruda:0")
print((x + x).cpu())

Prebuilt Windows wheels are available as artifacts of successful RUDA Torch Windows build runs. Select the wheel artifact, extract its .whl, and install that file with python -m pip install --no-deps. The wheel includes the native DLL and targets Windows x64, CPython 3.13 and PyTorch 2.13.0+cu130; install that matching PyTorch build first. Artifacts expire after seven days. ruda-torch-native is a source-build component, not a crates.io package.

Linux/Colab users can use precompiled bundles from GitHub Releases without compiling Rust or C++ locally. These are not pip wheels; use the matching source revision and the Python, PyTorch and glibc requirements recorded in the bundle manifest.

Stack Organization

One repository, multiple crates with clearly defined responsibilities. From domain libraries to higher-level frameworks, the stack is organized in layers and developed together.

Layer Components
Shared contracts ruda-core
Compilation and kernels ruda-compiler, ruda-kernel, macro components
Runtime and driver backends ruda, ruda-driver-cuda/cpu/wgpu/hip
Domain libraries ruBLAS, ruDNN, ruPRIM, ruFFT, ruRAND, ruSPARSE
Experimental numerical science ruSOLVER, ruINTEGRATE
Collective communication ruCCL, ruda-communication
Tensors and frameworks ruda-tensor*, ruda-autodiff, ruda-fusion
PyTorch integration ruda-torch-native (Rust), ruda_torch (Python)
Models and data ruda-model, ruda-nn, ruda-optim, ruda-store, ruda-dataset

Native GPU Inference

Paged attention requires contiguous FP32/FP16/BF16 tensors of the same dtype on the same device and execution queue. Forward and first-order backward are available, without arbitrary external masks or quantized KV caches. PyTorch defaults to atomic history gradients; backward_strategy="ordered" selects the atomic-free path. Ordered history-row caching and physical-page compaction are opt-in. MLA/MoE are reusable components; complete model adapters must supply projections, positional encoding, routing parameters and cache ownership.

Rust Training and Distributed Tensors

See Training and saving state and the ruCCL guide for API contracts and usage.

Experimental Numerical Science

rusolver adds host real/complex factorizations, SVD, sparse LU, row-partitioned CG and analytic pullbacks; opt-in FP32 batched LU/Cholesky/QR/eigen/CG device paths are separate. ruintegrate adds host quadrature, infinite-domain transforms, RK45, stiff BDF1 and event location. First-order host solver graph integration is opt-in via ruda-autodiff/solver-host. Extended scope.

See the guides for convergence and backend restrictions. The packages are workspace members but not default members.

Paths to Hardware

Explore and Contribute

If you care about Rust, GPU kernels, compilers, or high-performance computing, join us in taking this stack further and making it faster.

Origins and Licensing

Third-party notices

Original Ruda software code that the project has the right to license is available under the Apache License 2.0. Third-party files remain under their original licenses; the root license does not override the MIT OR Apache-2.0 declarations in migrated components.