Documentation / EnglishView source ↗

ruBLAS User Guide

Compute libraries · Runtime API · 中文 | 日本語 | Deutsch | Русский

1. Overview and features

ruBLAS provides linear algebra operations and device tensor interfaces.

Feature Module
tensor-vector rublas::tensor_vector
tensor-matmul rublas::tensor_matmul
tensor-matmul-autotune Tensor matrix multiplication autotuning
tensor-int4 rublas::tensor_int4
tensor-grouped rublas::tensor_grouped

The Cargo package is rublas. Select the required features and disable unnecessary defaults for general-purpose paths. See Cargo.toml.

2. Matrix multiplication, vectors, and INT4

Tensor matrix multiplication

rublas::tensor_matmul::matmul takes lhs, rhs, optional output out, MatmulStrategy, and output DType. It returns a device tensor or MatmulSetupError. Pass an existing output through out, or omit it to allocate one.

Strategies include Ruda, CmmaResidueFirst, Naive, and Autotune when tensor-matmul-autotune is enabled. The default is Ruda without that feature and Autotune with it. For quantized inputs, Naive may dequantize and compute after its initial path fails; this is not native INT4 computation.

CmmaResidueFirst explicitly selects the Tensor Core path that processes a partial K32 tile first, without requiring a materialized padded matrix. lhs and rhs must have matching unquantized BF16 or F16 dtype and compatible matrix dimensions. The device must support the path's matrix instructions. This function returns F32 output:

use ruda_core::tensor::DType;
use ruda_kernel::{dsl::Runtime, tensor::RudaTensor};
use rublas::{
    kernel_ir::definition::MatmulSetupError,
    tensor_matmul::{MatmulStrategy, matmul},
};

fn residue_matmul<R: Runtime>(
    lhs: RudaTensor<R>,
    rhs: RudaTensor<R>,
) -> Result<RudaTensor<R>, MatmulSetupError> {
    matmul(lhs, rhs, None, MatmulStrategy::CmmaResidueFirst, DType::F32)
}

Inputs [M, K] and [K, N] produce [M, N]. Invalid setup or missing device capabilities return MatmulSetupError rather than switching to Naive.

See the matrix multiplication entry point.

Vector cross product

rublas::tensor_vector::cross(lhs, rhs, dim) requires the selected dimension to have length 3 and returns a device tensor. Computing along a non-final dimension involves permutation and contiguous conversion. See cross.rs.

AWQ INT4

rublas::tensor_int4::AwqGemm::new takes qweight, qzeros, scales, optional bias, and group_size to construct packed weights. forward accepts F16 input and returns F16 output or Int4Error.

For input dimension K, output dimension N, and group size G:

Data Dtype Shape
qweight Packed I32 [K, N/8]
qzeros Packed I32 [K/G, N/8]
scales F16 [K/G, N]
bias (optional) F16 [N]

K, N, and G must be nonzero; K must be divisible by G and N by 8. The input's final dimension is K, replaced by N in the output. All operands must share a device. Packed bit order must match the AWQ kernel; an arbitrary INT4 file cannot be used directly as qweight.

See the INT4 interface for the object and its checks.

3. Grouped matrix multiplication

rublas::tensor_grouped::grouped_matmul_nt<R: Runtime> takes input, weights, and row_experts as RudaTensor<R> values. It returns Result<RudaTensor<R>, GroupedMatmulError>.

Argument Shape Requirements
input [M, K] Non-quantized F32, F16, or BF16
weights [E, N, K] Same dtype and device as input
row_experts [M] Non-quantized U32 on the same device
Output [M, N] Same dtype as input

Row m selects the weight matrix indexed by row_experts[m] and computes dot products between the input row and each row of that matrix. The final two weight dimensions participate in transposed form without requiring a materialized transpose.

4. Grouped multiplication semantics

Source: grouped interface and kernel.

5. Integration with other libraries

ruDNN MoE uses grouped multiplication for expert projections. Quantized weights use the separate tensor_int4 module; ordinary floating-point grouped multiplication is not INT4 expert computation.

grouped_matmul_nt uses scalar accumulation. Measure matrix multiplication strategies for your dtype, shape, and backend.

6. Segmented expert matrix multiplication

With tensor-grouped, rublas::tensor_grouped::grouped_matmul_nt_segmented(input, weights, row_experts, offsets, strategy) adds device-side expert segments to the existing grouped interface. Input is [M, K], weights [E, N, K], row experts U32 [M], and offsets U32 [E + 1]; output is [M, N] in the input dtype. Operands must be unquantized and share a device and execution queue.

This is an unsafe Rust API: offsets must be a nondecreasing exclusive prefix starting at 0 and ending at M, and every row in [offsets[e], offsets[e + 1]) must belong to expert e. row_experts must describe the same immutable dispatch. Shape checks do not validate these device values.

GroupedStrategy::Scalar uses the existing scalar kernel. TensorCore explicitly requires matching F16/BF16 inputs, supported 16×16×16 cooperative matrix operations, a 32-lane plane, sufficient shared memory and a legal launch grid; unsupported setup returns GroupedMatmulError. The kernel uses FP32 accumulation and handles partial tiles. Auto selects this path when supported and otherwise uses the scalar kernel; compilation, launch or numerical failures are not fallback conditions.

For a safe MoE entry point with internally constructed offsets, use rudnn::moe::SwiGluExperts::forward_dispatched_with_strategy. Its existing forward_dispatched entry point remains scalar by default.

7. Segmented backward

grouped_matmul_nt_backward_segmented(input, weights, grad_output, row_experts, offsets) returns GroupedBackward { dinput, dweights }. Input is [M, K], weights [E, N, K], and grad_output [M, N], all with matching F32/F16/BF16 dtype, device and queue. dinput retains the input dtype; dweights is FP32 with shape [E, N, K]. Empty expert segments produce zero weight gradients. Row IDs and offsets use the same U32 layouts and immutable prefix invariants as segmented forward; this remains an unsafe API without host readback of device metadata.

The default wrapper selects GroupedStrategy::Scalar. grouped_matmul_nt_backward_segmented_with_strategy(..., strategy) explicitly selects Scalar, Auto or TensorCore, independently of forward. The cooperative path uses 16×16×16 tiles, FP32 accumulation and FP32 weight gradients, and requires supported F16/BF16 hardware and launch limits. TensorCore errors if unsupported; Auto falls back only for unsupported capabilities, not compilation or execution errors. Inputs may be made contiguous. Different reduction orders need not be bitwise identical.

For a safe expert-level training path, use rudnn::moe::SwiGluExperts::forward_dispatched_training followed by ExpertTrainingCache::backward_with_strategy.