Installation and Quickstart
Documentation · Next: Programming guide · 中文 | 日本語 | Deutsch | Русский
1. Choose your entry point
| Task | Entry point |
|---|---|
| Write GPU kernels | ruda-kernel::dsl and a device runtime |
| Use matrix multiplication, FFTs, or reductions | Compute libraries |
| Work with tensors and frameworks | Tensors and frameworks |
| Train models and save state | Training guide |
| Load models and generate text or process images | Model inference guide |
| Integrate a device backend | Driver API |
Use Ruda from the source workspace.
2. Prepare an NVIDIA environment
You need Rust/Cargo, a linker toolchain for your platform, an NVIDIA GPU driver, and the CUDA Toolkit. The CUDA backend includes NVRTC and toolkit build dependencies; enabling direct PTX does not remove them.
Check your environment from the source root:
rustc --version --verbose
cargo --version
nvidia-smi
nvcc --version
cargo metadata --no-deps --format-version 1 --offline --locked
These commands do not compile Ruda. --offline requires the dependencies needed for resolution to be cached locally.
Set CUDA_PATH to select the CUDA Toolkit root. On Windows, point it to an installed version directory, not the parent containing multiple versions. See the CUDA installation path interface.
3. Build the example
With the build environment ready, run:
cargo build --locked -p ruda-driver-cuda --features direct-ptx --example ptx-runtime
The ptx-runtime example requires direct-ptx. Enabling this feature alone does not change the default compiler.
4. Select a compilation path and run
Choose one path in a separate PowerShell session. Resolve any command failure before continuing.
Default CUDA C++/NVRTC path:
$env:RUDA_CUDA_COMPILER = 'nvrtc'
cargo run --locked -p ruda-driver-cuda --features direct-ptx --example ptx-runtime
Direct PTX path:
$env:RUDA_CUDA_COMPILER = 'ptx'
$env:RUDA_PTX_VERSION = '8.0'
cargo run --locked -p ruda-driver-cuda --features direct-ptx --example ptx-runtime
The PTX version must match your target GPU and driver. See the PTX backend reference.
The example performs FP32 addition at several lengths, checks results and tail sentinels across repeated executions, and prints cache counters. See Examples and tutorials for optional cases.
5. Continue developing
Follow device selection, data upload, and kernel launch in the example, then read the programming guide. For build, driver loading, or execution errors, use Debugging and diagnostics to locate the failing stage.