Documentation / EnglishView source ↗

Backend selection and composition

Documentation · Tensor recipes · 中文

Kernel runtime versus tensor backend

cargo add ruda --features cuda selects the CUDA kernel facade. It does not turn ruda into the typed tensor, model, and optimizer crates. For Tensor<B, D>, select a tensor backend and enable ruda-tensor/api (or api-std).

Execution target Tensor backend Device Selection
CPU ruda_tensor_host::Host HostDevice ruda-tensor-host; std, optional simd/rayon
NVIDIA CUDA ruda_tensor_device::cuda::Cuda<f32, i32> CudaDevice ruda-tensor-device/cuda-default
AMD ROCm ruda_tensor_rocm::Rocm<f32, i32> RocmDevice ruda-tensor-rocm and ROCm/HIP runtime
WGPU ruda_tensor_wgpu::Wgpu<f32, i32> WgpuDevice Select vulkan, metal, or webgpu for the intended target
LibTorch ruda_tensor_tch::LibTorch LibTorchDevice Matching LibTorch installation
Multiple local backends ruda_tensor_router::Router<(B1, B2)> duo::MultiDevice<B1, B2> Explicit device variant
Remote execution ruda_tensor_remote::RemoteBackend RemoteDevice client feature and a running server

Hardware drivers and target toolchains are still required. Enabling a Cargo feature does not install them or automatically provide all backends. The WGPU path selects a graphics compute API rather than automatically switching between native CUDA and HIP.

Ascend is exposed in this repository through ruda-driver-cann, ruda-ascend-kernels, the compiler's Ascend paths, and rublas::cann. These are separate from the typed tensor backend choices above; see the compiler guide and ruBLAS guide.

CUDA tensor dependency

For typed tensor programs, add:

[dependencies]
ruda-tensor = { version = "0.21", default-features = false, features = ["api-std"] }
ruda-tensor-device = { version = "0.21", default-features = false, features = ["cuda-default"] }

Use ruda_tensor_device::cuda::{Cuda, CudaDevice} and choose Cuda<f32, i32> as the backend type. The full model/optimizer example is in Training.

cuda exposes the adapter; cuda-fusion wraps it in Fusion; cuda-default combines CUDA fusion with default runtime/backend facilities. For ROCm and WGPU, their fusion features determine whether the exported backend alias includes the fusion wrapper. Cargo features are additive: another dependency enabling fusion can affect the final alias.

Route tensors between CPU and WGPU

This example explicitly uses Vulkan, requiring a compatible Vulkan adapter. Create a binary project with:

[dependencies]
ruda-tensor = { version = "0.21", default-features = false, features = ["api-std"] }
ruda-tensor-host = { version = "0.21", default-features = false, features = ["std"] }
ruda-tensor-router = "0.21"
ruda-tensor-wgpu = { version = "0.21", features = ["vulkan"] }

Use this src/main.rs and run cargo run:

use ruda_tensor::{api::Tensor, DType};
use ruda_tensor_host::{Host, HostDevice};
use ruda_tensor_router::{duo::MultiDevice, Router};
use ruda_tensor_wgpu::{Wgpu, WgpuDevice};

type B = Router<(Host, Wgpu<f32, i32>)>;
type Device = MultiDevice<Host, Wgpu<f32, i32>>;

fn main() {
    let cpu = Device::B1(HostDevice);
    let gpu = Device::B2(WgpuDevice::default());
    let x = Tensor::<B, 1>::from_data([1.0f32, 2.0, 3.0], (&cpu, DType::F32));
    let x_gpu = x.to_device(&gpu);
    let result = (x_gpu.clone() + x_gpu).to_device(&cpu).into_data();
    assert_eq!(result.as_slice::<f32>().unwrap(), &[2.0, 4.0, 6.0]);
}

B1 and B2 follow tuple order. The router's default device selects the first backend; explicit variants make placement unambiguous. trio and quad provide corresponding three- and four-backend device enums.

Router uses DirectByteChannel and ByteBridge. A cross-backend transfer goes through TensorData; it is not a zero-copy or GPU peer-to-peer guarantee. to_device moves the tensor explicitly. The router does not automatically choose the fastest backend, split a model, or replace unsupported operations with a CPU implementation.

Remote tensor execution

The server selects the actual backend; the client sends tensor operations over WebSocket. A CPU server is sufficient for learning the protocol. Use these dependencies in a new project:

[dependencies]
ruda-tensor = { version = "0.21", default-features = false, features = ["api-std"] }
ruda-tensor-host = { version = "0.21", default-features = false, features = ["std"] }
ruda-tensor-remote = { version = "0.21", default-features = false, features = ["client", "server"] }

Create src/bin/server.rs:

use ruda_tensor_host::{Host, HostDevice};

fn main() {
    ruda_tensor_remote::server::start_websocket::<Host>(HostDevice, 3000);
}

Create src/bin/client.rs:

use ruda_tensor::{api::Tensor, DType};
use ruda_tensor_remote::{RemoteBackend, RemoteDevice};

fn main() {
    let device = RemoteDevice::new("ws://127.0.0.1:3000");
    let x = Tensor::<RemoteBackend, 1>::from_data(
        [1.0f32, 2.0, 3.0], (&device, DType::F32),
    );
    let result = (x.clone() + x).into_data();
    assert_eq!(result.as_slice::<f32>().unwrap(), &[2.0, 4.0, 6.0]);
}

Run cargo run --bin server, then run cargo run --bin client in a second terminal. The convenience server binds 0.0.0.0:3000, not only loopback; its built-in transport does not configure authentication or TLS. The client URL above is for a same-machine connection.

start_websocket owns its Tokio runtime and blocks until the server stops. Inside an existing async runtime, use start_websocket_async::<B>(device, port).await. The server backend must implement BackendIr; replacing the server's backend also requires its dependencies and hardware. Remote execution is distinct from ruCCL distributed collectives.

Choosing wrappers and boundaries

API entry points