Ruda Documentation
Documentation index · 中文 | 日本語 | Deutsch | Русский
From your first GPU kernel to compute libraries, tensor training, and local model inference. Start with your task, then explore the programming guides and API references.
Start here
| Your task | Reading path |
|---|---|
| Run your first GPU kernel | Installation and quickstart → Programming guide |
| Use matrix, sparse, or neural network operations | Compute libraries → Tensors and frameworks |
| Train models, accumulate gradients, and save state | Training and saving state |
| Fine-tune a local model with LoRA or NF4 | Fine-tuning and recovery |
| Load local models, generate text, or process images | Model loading and inference |
Getting started
- Installation and quickstart: source setup, backend selection, and your first example.
- Examples and tutorials: vector addition, tensors, shared memory, and half precision.
Programming guides
- Ruda programming guide: host and device code, execution hierarchy, memory, synchronization, and safety.
- Tensors and frameworks: device tensors, library dispatch, fusion, and automatic differentiation.
- Tensor recipes: complete CPU examples for matrix multiplication, layouts, dtype selection, and gradients.
- Backend selection and composition: CUDA/ROCm/WGPU selection, local routing, and remote execution.
- Data pipelines and model storage: batching samples, saving weights, and importing checkpoint formats.
- Training and saving state: training steps, FP32 masters/accumulation, mixed-storage checkpoints, token-weighted replicas, differentiable collectives and learning-rate schedules.
- Ranks, devices and distributed training: explicit rendezvous, device placement, collective order, weighted gradients and rank-local recovery.
- Architecture components and Python Muon: mHC, DSA/CSA/HCA, compressed KV caches and hybrid model composition.
- LoRA and NF4 fine-tuning: exact target selection, streamed weights, causal supervision, token-weighted accumulation, adapters and restart checkpoints.
- General PyTorch model compiler: AOT forward/backward, native partitions, options and cache ownership.
- Fixed-address PyTorch subgraphs: explicit GraphOp plans, output lifetime, scratch reuse and first-order training.
- Model loading and inference: ruLLM, text and image inputs, sampling, AWQ, and continuous batching.
Compilation and execution
- Compiler guide: the Rust kernel frontend, IR, CUDA C++/NVRTC, and direct PTX.
- PTX backend reference: target configuration, compilation output, constraints, and errors.
- Shared stack autotuning: participating operators, offline calibration, timing, cache identity and policy parameters.
API references
- Full-stack crate and API index: package names, Rust imports, responsibilities, and source entry points across the workspace.
- Runtime API: device clients, memory, submission, readback, and synchronization.
- Driver API and backends: backend types, device selection, and runtime integration.
- Native PyTorch API: device/component versions, shape/dtype contracts, normalization, optimizers, streams, attention, sequence training and quantization.
- Compute library reference: library selection, Cargo features, and entry points.
Compute libraries
| Library | Guide |
|---|---|
| ruBLAS | Linear algebra and grouped matrix multiplication |
| ruDNN | Neural network operations and MoE |
| ruPRIM | Reductions, scans, and indexing |
| ruFFT | Fast Fourier transforms |
| ruRAND | Random number generation |
| ruSPARSE | Sparse computation |
| ruCCL | Collective communication |
Debugging and compatibility
- Debugging and diagnostics: compilation errors, asynchronous errors, caching, and numerical checks.
- Compatibility guide: CUDA concepts, backend differences, and API and compilation boundaries.
- Contributing: reporting issues and development conventions.
For a first project, follow the quickstart, programming guide, and relevant library guide. Consult the compiler and API references when developing backends or kernels.