Documentation / EnglishView source ↗

Full-stack crate and API index

Documentation · 中文

Start at the right layer

Task First API Guide
Write a GPU kernel ruda, kernel DSL, runtime client Programming guide
Use typed tensors and gradients ruda_tensor::api::Tensor, Autodiff<B> Tensor recipes
Choose a tensor execution target Host, Cuda, Rocm, Wgpu, Router, Remote Backend composition
Call a domain operator ruBLAS, ruDNN, ruPRIM, and other domain crates Compute libraries
Build and train a model Module, ruda-nn, ruda-optim Training
Train with half storage and FP32 updates Module::to_dtype, Fp32MasterOptimizer, GradientsAccumulator::accumulate_with_dtype Mixed-precision training
Preserve mixed storage and pending gradients TrainingRecord::capture_with_dtypes, restore_with_dtypes Training state
Synchronize replicas or differentiate collectives DataParallel, ruda_autodiff::collective Distributed training
Load samples or weights Dataset, DataLoaderBuilder, ModuleSnapshot Data and storage
Run local model inference rullm Model inference

Package names and source entry points

The Cargo package name, repository directory, and Rust import name can differ. For example, ruFFT lives in ruFFT/, is installed as ruda-fft, and is imported as rufft; ruSOLVER is installed as ruda-solver and imported as rusolver.

The tables cover the resolved workspace, including the path-dependent CANN driver and the test-support packages. Package links open their manifests; source links open the crate entry point. Features and versions are specified by each manifest, not by the directory name. Test fixtures and the native PyTorch extension are not ordinary crates.io installation targets.

Kernel, compiler, and runtime

Cargo package Rust import / source Responsibility
ruda ruda Kernel facade and runtime re-exports; start here for GPU kernels
ruda-core ruda_core Shared IR, tensor data, dtype, device, and memory contracts
ruda-compiler ruda_compiler Kernel frontends and target code generation
ruda-kernel ruda_kernel Rust kernel DSL, launch interfaces, and device tensor utilities
ruda-runtime ruda_runtime Compute clients/servers, storage, streams, and scheduling
ruda-kernel-macros ruda_kernel_macros Procedural macros translating kernel Rust to IR
ruda-ir-macros ruda_ir_macros Derive helpers for IR implementations

Drivers and Ascend programs

Cargo package Rust import / source Responsibility
ruda-driver-cuda ruda_driver_cuda CUDA devices, allocation, compilation, and launch
ruda-driver-hip ruda_driver_hip AMD HIP runtime adapter
ruda-hip-sys ruda_hip_sys Low-level HIP runtime bindings
ruda-driver-wgpu ruda_driver_wgpu WGPU setup, storage, and compute execution
ruda-driver-cpu ruda_driver_cpu CPU kernel runtime driver, separate from Host tensor backend
ruda-driver-cann ruda_driver_cann Dynamically loaded CANN AscendCL interfaces
ruda-ascend-kernels ruda_ascend_kernels Rust Ascend device programs and instruction lowering

Compute libraries

Cargo package Rust import / source Responsibility
rublas rublas Matrix/vector operations and backend dispatch
ruDNN rudnn Attention, convolution, pooling, normalization, and MoE
ruPRIM ruprim Reductions, scans, indexing, and elementwise kernels
ruda-fft rufft Fourier transforms
ruRAND rurand Random sampling and distributions
ruSPARSE rusparse Sparse formats and operators
ruTENSOR rutensor Tensor contractions, permutations, and reductions
ruCCL ruccl Collective communication algorithms
ruda-solver rusolver Host scientific solvers and opt-in device solve kernels
ruintegrate ruintegrate Host quadrature, ODE integration, and event location
rublas-host rublas_host CPU strided and batched matrix multiplication
ruDNN-host rudnn_host CPU neural-network operator implementations
ruPRIM-host ruprim_host CPU tensor primitives and indexing
ruFFT-host rufft_host CPU real Fourier transforms
ruRAND-host rurand_host Host random-number generation

Tensors, differentiation, and execution composition

Cargo package Rust import / source Responsibility
ruda-tensor ruda_tensor Backend contracts and api::Tensor behind the api feature
ruda-tensor-config ruda_tensor_config Shared autodiff/fusion configuration
ruda-tensor-device ruda_tensor_device DeviceBackend, CUDA adapter, and domain-library dispatch
ruda-tensor-host ruda_tensor_host Host CPU tensor backend and strided layouts
ruda-tensor-wgpu ruda_tensor_wgpu WGPU tensor backend and device setup
ruda-tensor-rocm ruda_tensor_rocm ROCm tensor adapter
ruda-tensor-tch ruda_tensor_tch LibTorch tensor adapter
ruda-autodiff ruda_autodiff Gradient graph, backward pass, and recomputation
ruda-fusion ruda_fusion Tensor operation fusion planning and execution
ruda-tensor-router ruda_tensor_router Local multi-backend routing and byte bridges
ruda-tensor-remote ruda_tensor_remote Remote tensor client/server
ruda-communication ruda_communication Transport protocols, WebSocket, and tensor data service

Models, training, storage, and integrations

Cargo package Rust import / source Responsibility
ruda-model ruda_model Module parameters, configuration, records, and data loaders
ruda-model-macros ruda_model_macros Config, Module, and Record derive macros
ruda-model-codegen ruda_model_codegen Code generation used by the model derive macros
ruda-nn ruda_nn Neural-network layers, activations, and losses
ruda-optim ruda_optim Optimizers, gradient accumulation/clipping, and schedules
ruda-dataset ruda_dataset Indexed datasets, sources, and transforms
ruda-io ruda_io Host I/O and optional network downloads
ruda-store ruda_store Model snapshots, Rudapack, safetensors, and PyTorch import
ruda-llm rullm Model loading and autoregressive inference
ruda-torch-native ruda_torch_native Native cdylib for the Python ruda_torch package; source-build component

Test and consumer support

Cargo package Rust import / source Responsibility
ruda-test-runtime ruda_test_runtime Runtime/kernel test infrastructure
ruda-test-utils ruda_test_utils Kernel test helpers
ruda-facade-consumer ruda_facade_consumer External-consumer fixture for facade feature wiring
ruda-store-pytorch-tests ruda_store_pytorch_tests PyTorch format interoperability tests
ruda-store-safetensors-tests ruda_store_safetensors_tests Safetensors format interoperability tests

Find an individual method

Native Python methods are documented in the PyTorch API reference, model compiler, static graphs and LoRA/NF4 fine-tuning guide. Shared operator/application selection uses the stack autotuning policy.

The architecture guide documents mHC, compressed attention/caches and Python Muon. The distributed training guide covers explicit devices, rendezvous, replica initialization, weighted reduction and rank-local recovery.

Use the typed tensor methods under ruda-tensor/src/api, backend traits under ruda-tensor/src/backend, and domain-specific modules linked from each library guide. Enable the feature that exposes the module before using its symbols.

For device memory, submission, and synchronization, read the Runtime API. For backend-specific initialization and launch contracts, read the Driver API. For model parameter and record types, start with ruda-model exports, rather than the compiler's similarly named IR types.