Documentation / EnglishView source ↗

Installation and Quickstart

Documentation · Next: Programming guide · 中文 | 日本語 | Deutsch | Русский

1. Choose your entry point

Task Entry point
Write GPU kernels ruda-kernel::dsl and a device runtime
Use matrix multiplication, FFTs, or reductions Compute libraries
Work with tensors and frameworks Tensors and frameworks
Train models and save state Training guide
Load models and generate text or process images Model inference guide
Integrate a device backend Driver API

Use Ruda from the source workspace.

2. Prepare an NVIDIA environment

You need Rust/Cargo, a linker toolchain for your platform, an NVIDIA GPU driver, and the CUDA Toolkit. The CUDA backend includes NVRTC and toolkit build dependencies; enabling direct PTX does not remove them.

Check your environment from the source root:

rustc --version --verbose
cargo --version
nvidia-smi
nvcc --version
cargo metadata --no-deps --format-version 1 --offline --locked

These commands do not compile Ruda. --offline requires the dependencies needed for resolution to be cached locally.

Set CUDA_PATH to select the CUDA Toolkit root. On Windows, point it to an installed version directory, not the parent containing multiple versions. See the CUDA installation path interface.

3. Build the example

With the build environment ready, run:

cargo build --locked -p ruda-driver-cuda --features direct-ptx --example ptx-runtime

The ptx-runtime example requires direct-ptx. Enabling this feature alone does not change the default compiler.

4. Select a compilation path and run

Choose one path in a separate PowerShell session. Resolve any command failure before continuing.

Default CUDA C++/NVRTC path:

$env:RUDA_CUDA_COMPILER = 'nvrtc'
cargo run --locked -p ruda-driver-cuda --features direct-ptx --example ptx-runtime

Direct PTX path:

$env:RUDA_CUDA_COMPILER = 'ptx'
$env:RUDA_PTX_VERSION = '8.0'
cargo run --locked -p ruda-driver-cuda --features direct-ptx --example ptx-runtime

The PTX version must match your target GPU and driver. See the PTX backend reference.

The example performs FP32 addition at several lengths, checks results and tail sentinels across repeated executions, and prints cache counters. See Examples and tutorials for optional cases.

5. Continue developing

Follow device selection, data upload, and kernel launch in the example, then read the programming guide. For build, driver loading, or execution errors, use Debugging and diagnostics to locate the failing stage.