ruBLAS User Guide
Compute libraries · Runtime API · 中文 | 日本語 | Deutsch | Русский
1. Overview and features
ruBLAS provides linear algebra operations and device tensor interfaces.
| Feature | Module |
|---|---|
tensor-vector |
rublas::tensor_vector |
tensor-matmul |
rublas::tensor_matmul |
tensor-matmul-autotune |
Tensor matrix multiplication autotuning |
tensor-int4 |
rublas::tensor_int4 |
tensor-grouped |
rublas::tensor_grouped |
The Cargo package is rublas. Select the required features and disable unnecessary defaults for general-purpose paths. See Cargo.toml.
2. Matrix multiplication, vectors, and INT4
Tensor matrix multiplication
rublas::tensor_matmul::matmul takes lhs, rhs, optional output out, MatmulStrategy, and output DType. It returns a device tensor or MatmulSetupError. Pass an existing output through out, or omit it to allocate one.
Strategies include Ruda, CmmaResidueFirst, Naive, and Autotune when tensor-matmul-autotune is enabled. The default is Ruda without that feature and Autotune with it. For quantized inputs, Naive may dequantize and compute after its initial path fails; this is not native INT4 computation.
CmmaResidueFirst explicitly selects the Tensor Core path that processes a partial K32 tile first, without requiring a materialized padded matrix. lhs and rhs must have matching unquantized BF16 or F16 dtype and compatible matrix dimensions. The device must support the path's matrix instructions. This function returns F32 output:
use ruda_core::tensor::DType;
use ruda_kernel::{dsl::Runtime, tensor::RudaTensor};
use rublas::{
kernel_ir::definition::MatmulSetupError,
tensor_matmul::{MatmulStrategy, matmul},
};
fn residue_matmul<R: Runtime>(
lhs: RudaTensor<R>,
rhs: RudaTensor<R>,
) -> Result<RudaTensor<R>, MatmulSetupError> {
matmul(lhs, rhs, None, MatmulStrategy::CmmaResidueFirst, DType::F32)
}
Inputs [M, K] and [K, N] produce [M, N]. Invalid setup or missing device capabilities return MatmulSetupError rather than switching to Naive.
See the matrix multiplication entry point.
Vector cross product
rublas::tensor_vector::cross(lhs, rhs, dim) requires the selected dimension to have length 3 and returns a device tensor. Computing along a non-final dimension involves permutation and contiguous conversion. See cross.rs.
AWQ INT4
rublas::tensor_int4::AwqGemm::new takes qweight, qzeros, scales, optional bias, and group_size to construct packed weights. forward accepts F16 input and returns F16 output or Int4Error.
For input dimension K, output dimension N, and group size G:
| Data | Dtype | Shape |
|---|---|---|
| qweight | Packed I32 | [K, N/8] |
| qzeros | Packed I32 | [K/G, N/8] |
| scales | F16 | [K/G, N] |
| bias (optional) | F16 | [N] |
K, N, and G must be nonzero; K must be divisible by G and N by 8. The input's final dimension is K, replaced by N in the output. All operands must share a device. Packed bit order must match the AWQ kernel; an arbitrary INT4 file cannot be used directly as qweight.
See the INT4 interface for the object and its checks.
3. Grouped matrix multiplication
rublas::tensor_grouped::grouped_matmul_nt<R: Runtime> takes input, weights, and row_experts as RudaTensor<R> values. It returns Result<RudaTensor<R>, GroupedMatmulError>.
| Argument | Shape | Requirements |
|---|---|---|
| input | [M, K] | Non-quantized F32, F16, or BF16 |
| weights | [E, N, K] | Same dtype and device as input |
| row_experts | [M] | Non-quantized U32 on the same device |
| Output | [M, N] | Same dtype as input |
Row m selects the weight matrix indexed by row_experts[m] and computes dot products between the input row and each row of that matrix. The final two weight dimensions participate in transposed form without requiring a materialized transpose.
4. Grouped multiplication semantics
- K, N, and E must be positive; M may be zero.
- Expert indices outside the valid range indicate padding and produce zero output rows.
- The kernel accumulates in FP32, then casts to the input dtype.
- The entry point makes input and weights contiguous when necessary, which may copy data.
- Relevant element counts must fit U32 indexing.
- Invalid arguments return
GroupedMatmulError. Handle asynchronous execution errors during readback or synchronization.
Source: grouped interface and kernel.
5. Integration with other libraries
ruDNN MoE uses grouped multiplication for expert projections. Quantized weights use the separate tensor_int4 module; ordinary floating-point grouped multiplication is not INT4 expert computation.
grouped_matmul_nt uses scalar accumulation. Measure matrix multiplication strategies for your dtype, shape, and backend.
6. Segmented expert matrix multiplication
With tensor-grouped, rublas::tensor_grouped::grouped_matmul_nt_segmented(input, weights, row_experts, offsets, strategy) adds device-side expert segments to the existing grouped interface. Input is [M, K], weights [E, N, K], row experts U32 [M], and offsets U32 [E + 1]; output is [M, N] in the input dtype. Operands must be unquantized and share a device and execution queue.
This is an unsafe Rust API: offsets must be a nondecreasing exclusive prefix starting at 0 and ending at M, and every row in [offsets[e], offsets[e + 1]) must belong to expert e. row_experts must describe the same immutable dispatch. Shape checks do not validate these device values.
GroupedStrategy::Scalar uses the existing scalar kernel. TensorCore explicitly requires matching F16/BF16 inputs, supported 16×16×16 cooperative matrix operations, a 32-lane plane, sufficient shared memory and a legal launch grid; unsupported setup returns GroupedMatmulError. The kernel uses FP32 accumulation and handles partial tiles. Auto selects this path when supported and otherwise uses the scalar kernel; compilation, launch or numerical failures are not fallback conditions.
For a safe MoE entry point with internally constructed offsets, use rudnn::moe::SwiGluExperts::forward_dispatched_with_strategy. Its existing forward_dispatched entry point remains scalar by default.
7. Segmented backward
grouped_matmul_nt_backward_segmented(input, weights, grad_output, row_experts, offsets) returns GroupedBackward { dinput, dweights }. Input is [M, K], weights [E, N, K], and grad_output [M, N], all with matching F32/F16/BF16 dtype, device and queue. dinput retains the input dtype; dweights is FP32 with shape [E, N, K]. Empty expert segments produce zero weight gradients. Row IDs and offsets use the same U32 layouts and immutable prefix invariants as segmented forward; this remains an unsafe API without host readback of device metadata.
The default wrapper selects GroupedStrategy::Scalar. grouped_matmul_nt_backward_segmented_with_strategy(..., strategy) explicitly selects Scalar, Auto or TensorCore, independently of forward. The cooperative path uses 16×16×16 tiles, FP32 accumulation and FP32 weight gradients, and requires supported F16/BF16 hardware and launch limits. TensorCore errors if unsupported; Auto falls back only for unsupported capabilities, not compilation or execution errors. Inputs may be made contiguous. Different reduction orders need not be bitwise identical.
For a safe expert-level training path, use rudnn::moe::SwiGluExperts::forward_dispatched_training followed by ExpertTrainingCache::backward_with_strategy.