Documentation / EnglishView source ↗ Full-stack crate and API index
Documentation · 中文
Start at the right layer
| Task |
First API |
Guide |
| Write a GPU kernel |
ruda, kernel DSL, runtime client |
Programming guide |
| Use typed tensors and gradients |
ruda_tensor::api::Tensor, Autodiff<B> |
Tensor recipes |
| Choose a tensor execution target |
Host, Cuda, Rocm, Wgpu, Router, Remote |
Backend composition |
| Call a domain operator |
ruBLAS, ruDNN, ruPRIM, and other domain crates |
Compute libraries |
| Build and train a model |
Module, ruda-nn, ruda-optim |
Training |
| Train with half storage and FP32 updates |
Module::to_dtype, Fp32MasterOptimizer, GradientsAccumulator::accumulate_with_dtype |
Mixed-precision training |
| Preserve mixed storage and pending gradients |
TrainingRecord::capture_with_dtypes, restore_with_dtypes |
Training state |
| Synchronize replicas or differentiate collectives |
DataParallel, ruda_autodiff::collective |
Distributed training |
| Load samples or weights |
Dataset, DataLoaderBuilder, ModuleSnapshot |
Data and storage |
| Run local model inference |
rullm |
Model inference |
Package names and source entry points
The Cargo package name, repository directory, and Rust import name can differ. For example, ruFFT lives in ruFFT/, is installed as ruda-fft, and is imported as rufft; ruSOLVER is installed as ruda-solver and imported as rusolver.
The tables cover the resolved workspace, including the path-dependent CANN driver and the test-support packages. Package links open their manifests; source links open the crate entry point. Features and versions are specified by each manifest, not by the directory name. Test fixtures and the native PyTorch extension are not ordinary crates.io installation targets.
Kernel, compiler, and runtime
Drivers and Ascend programs
Compute libraries
| Cargo package |
Rust import / source |
Responsibility |
rublas |
rublas |
Matrix/vector operations and backend dispatch |
ruDNN |
rudnn |
Attention, convolution, pooling, normalization, and MoE |
ruPRIM |
ruprim |
Reductions, scans, indexing, and elementwise kernels |
ruda-fft |
rufft |
Fourier transforms |
ruRAND |
rurand |
Random sampling and distributions |
ruSPARSE |
rusparse |
Sparse formats and operators |
ruTENSOR |
rutensor |
Tensor contractions, permutations, and reductions |
ruCCL |
ruccl |
Collective communication algorithms |
ruda-solver |
rusolver |
Host scientific solvers and opt-in device solve kernels |
ruintegrate |
ruintegrate |
Host quadrature, ODE integration, and event location |
rublas-host |
rublas_host |
CPU strided and batched matrix multiplication |
ruDNN-host |
rudnn_host |
CPU neural-network operator implementations |
ruPRIM-host |
ruprim_host |
CPU tensor primitives and indexing |
ruFFT-host |
rufft_host |
CPU real Fourier transforms |
ruRAND-host |
rurand_host |
Host random-number generation |
Tensors, differentiation, and execution composition
Models, training, storage, and integrations
| Cargo package |
Rust import / source |
Responsibility |
ruda-model |
ruda_model |
Module parameters, configuration, records, and data loaders |
ruda-model-macros |
ruda_model_macros |
Config, Module, and Record derive macros |
ruda-model-codegen |
ruda_model_codegen |
Code generation used by the model derive macros |
ruda-nn |
ruda_nn |
Neural-network layers, activations, and losses |
ruda-optim |
ruda_optim |
Optimizers, gradient accumulation/clipping, and schedules |
ruda-dataset |
ruda_dataset |
Indexed datasets, sources, and transforms |
ruda-io |
ruda_io |
Host I/O and optional network downloads |
ruda-store |
ruda_store |
Model snapshots, Rudapack, safetensors, and PyTorch import |
ruda-llm |
rullm |
Model loading and autoregressive inference |
ruda-torch-native |
ruda_torch_native |
Native cdylib for the Python ruda_torch package; source-build component |
Test and consumer support
Find an individual method
Native Python methods are documented in the PyTorch API reference, model compiler, static graphs and LoRA/NF4 fine-tuning guide. Shared operator/application selection uses the stack autotuning policy.
The architecture guide documents mHC, compressed attention/caches and Python Muon. The distributed training guide covers explicit devices, rendezvous, replica initialization, weighted reduction and rank-local recovery.
Use the typed tensor methods under ruda-tensor/src/api, backend traits under ruda-tensor/src/backend, and domain-specific modules linked from each library guide. Enable the feature that exposes the module before using its symbols.
For device memory, submission, and synchronization, read the Runtime API. For backend-specific initialization and launch contracts, read the Driver API. For model parameter and record types, start with ruda-model exports, rather than the compiler's similarly named IR types.