cheatah
Module

fixarray

Benchmarks
Every measured table for fixarray — with its stamp (commit, host, harness) and the commands that reproduce it — is on the fixarray benchmarks page.

Fixed-extent vectors and matrices — the same mathematics as an ndarray::NDArray, with the shape moved into the type, for the small-and-hot regime (a renderer's transforms, a physics solver's contact frames, a filter's small state). Header-only; import fixarray (it auto-links ndarray for its element concepts). The heavy shape-generic numerics on the dynamic NDArray (LU/QR/SVD/eigen) live in the sibling linalg module.

The type

An NDArray carries its shape at runtime and its elements on the heap. That is what makes it general — and, for a 3-D direction or a 4×4 transform, it is the whole cost: an allocation and an indirection to move sixteen floats.

fixarray::Fixed<T, Dims...> is the same idea with the shape in the type. The elements live inline, so the value is trivially copyable and nothing allocates; the loops have compile-time trip counts and auto-vectorize. These are the types for high-performance work — transforms, contact frames, small filter state — where the same matrix is built and consumed millions of times a second. Everything else matches NDArray: the element types, the (row, column) index, numpy's vocabulary, and the answers.

using namespace cheatah::fixarray;
vec3f up{0.0F, 1.0F, 0.0F};          // 12 bytes, no allocation
mat4f mvp = projection * view;       // 64 bytes — exactly a push constant
vec3f n = normalize(cross(up, dir));

Aliases: vec2f…vec4f, vec2d…vec4d, mat2f…mat4f, mat2d…mat4d; Vec<T,N> and Mat<T,R,C> for anything else. Rank 1 (vector) and rank 2 (matrix).

From cheatah

import fixarray and use it from a .purr program — construct with the call form and operate with the free functions:

import io
import fixarray

fn length2(v: fixarray.Fixed<f32, 3>) -> f32 {   # module-qualified param/return
    return fixarray.dot(v, v)
}

fn main() {
    let a = fixarray.vec3f(1.0, 2.0, 3.0)        # construct (deduced type)
    let b: fixarray.vec3f = fixarray.vec3f(4.0, 5.0, 6.0)   # or annotate the type
    io.print(fixarray.dot(a, b))                 # 32
    let bytes: fixarray.Vec<u8, 3> = fixarray.Vec<u8, 3>(1, 2, 3)  # narrow element (1 byte each)
    io.print(length2(a))                         # 14
}
main()

Aliases (vec3f, mat4f, …) and the explicit fixarray.Fixed<T, Dims…> / fixarray.Vec<T, N> spellings all work as let/field/parameter/return annotations. A parameter of a fixarray type binds by reference, so pass a named vector (an lvalue) to your own functions. import fixarray alone pulls in ndarray transitively.

A matrix is stored column-major, unlike NDArray. The indexing you write does not change, but data() hands back columns. That is deliberate and measured: m * v becomes a sum of scaled columns — contiguous and vertical — instead of four horizontal dot products behind a shuffle network, and the buffer is already in the order GPU APIs expect, so uploading a transform is a copy rather than a transpose.

Speed

Benchmarked against GLM across the complete overlap of the two APIs — 160 pairs: every vector and matrix operation (arithmetic, dot, cross, length, normalize, distance, reflect, min/max/clamp, mix, step/smoothstep, matmul, transpose, determinant, inverse, matrixCompMult, outerProduct, inverseTranspose), at sizes 2/3/4 — the GLSL common functions and the extra matrix builtins at 3 and 4 — in both float and double (fixed_glm_bench.cpp). Both sides compile in the same translation unit with the same flags, so neither gets an instruction set the other lacks — and the benchmark verifies the two produce identical results before it times either, so a fast wrong answer can never masquerade as a win.

Fixed is faster than or at parity with GLM across the whole overlap, and slower on none. The generated tables that carry the exact tally — a representative sample, the complete comparison, and the reading of the parity floor that decides each verdict — are on the fixarray benchmarks page, and the bench_gate.sh hard gate keeps the claim true: it fails the build if any pair regresses past GLM (tolerant by ratio and an absolute floor, with a confirmation re-run so sub-nanosecond noise never flakes it).

Choices that earn it, each a comment in fixarray.hpp: a matrix is stored column-major, so m * v is a sum of contiguous columns rather than horizontal dot products behind a shuffle network; dot sums pairwise and, at width ≥ 4, packs the products into one SIMD multiply (also lowering the rounding error to O(log n)); normalize takes one reciprocal then multiplies; min/max/abs/clamp are branchless always-writes that lower to minps/maxps; and every elementwise op builds its result in one pass through from_indices (no default-zero then overwrite). No intrinsics — the code is shaped so the compiler vectorizes it, no library to link at all.

Classes

Functions

fn T dot(const Vec< T, N > &a, const Vec< T, N > &b) source#

Inner product of two vectors — Σ aᵢbᵢ, the same quantity dot(const NDArray&, const NDArray&) computes, without the allocation.

Summed pairwise, so it is both faster and more accurate than a left-to-right accumulation.

Template parameters
T

the element type.

N

the length.

Parameters
a, b

the vectors.

Returns

the inner product.

Complexity

O(N).

Allocation

none.

fn Vec< T, 3 > cross(const Vec< T, 3 > &a, const Vec< T, 3 > &b) source#

Cross product of two 3-vectors — the vector perpendicular to both, right-handed.

Template parameters
T

the element type.

Parameters
a, b

the vectors.

Returns

a × b.

Complexity

O(1).

Allocation

none.

fn T squared_norm(const Vec< T, N > &v) source#

The squared Euclidean length of a vector — dot(v, v).

Prefer it to norm when only comparing lengths: it avoids the square root.

Template parameters
T

the element type.

N

the length.

Parameters
v

the vector.

Returns

Σ vᵢ².

Complexity

O(N).

Allocation

none.

Warning

for a complex element this is the bilinear Σ vᵢ² (matching dot), not the Hermitian Σ |vᵢ|² — it is the squared Euclidean LENGTH only for real elements.

fn T norm(const Vec< T, N > &v) source#

The Euclidean length of a vector.

Template parameters
T

the element type; floating-point, since the result is a root.

N

the length.

Parameters
v

the vector.

Returns

sqrt(Σ vᵢ²).

Complexity

O(N).

Allocation

none.

fn Vec< T, N > normalize(const Vec< T, N > &v) source#

The unit vector pointing the same way as v.

Template parameters
T

the element type; floating-point.

N

the length.

Parameters
v

the vector; must not be the zero vector.

Returns

v / norm(v).

Parameters
std::domain_error

when v has zero length, since it has no direction.

Complexity

O(N).

Allocation

none.

fn Mat< T, C, R > transpose(const Mat< T, R, C > &m) source#

The transpose of a matrix — rows become columns.

Template parameters
T

the element type.

R

the rows.

C

the columns.

Parameters
m

the matrix.

Returns

the C×R transpose.

Complexity

O(R·C).

Allocation

none.

fn T trace(const Mat< T, N, N > &m) source#

The trace of a square matrix — the sum of its diagonal.

Template parameters
T

the element type.

N

the dimension.

Parameters
m

the matrix.

Returns

Σ mᵢᵢ.

Complexity

O(N).

Allocation

none.

fn Mat< T, R, C > matmul(const Mat< T, R, K > &a, const Mat< T, K, C > &b) source#

Matrix product — the same A·B matmul(const NDArray&, const NDArray&) computes, with the shapes checked by the compiler rather than at runtime.

Template parameters
T

the element type.

R

the rows of a.

K

the shared dimension.

C

the columns of b.

Parameters
a, b

the matrices.

Returns

the R×C product.

Complexity

O(R·K·C).

Allocation

none.

fn operator* · 2 overloads
Mat< T, R, C > operator*(const Mat< T, R, K > &a, const Mat< T, K, C > &b)source#
Vec< T, R > operator*(const Mat< T, R, C > &m, const Vec< T, C > &v)source#

Matrix product, spelled a * b.

Template parameters
T

the element type.

R

the rows of a.

K

the shared dimension.

C

the columns of b.

Parameters
a, b

the matrices.

Returns

matmul(a, b).

Complexity

O(R·K·C).

Allocation

none.

fn determinant · 3 overloads
T determinant(const Mat< T, 2, 2 > &m)source#
T determinant(const Mat< T, 3, 3 > &m)source#
T determinant(const Mat< T, 4, 4 > &m)source#

The determinant of a 2×2 matrix, in closed form.

Template parameters
T

the element type.

Parameters
m

the matrix.

Returns

det(m).

Complexity

O(1).

Allocation

none.

fn inverse · 3 overloads
Mat< T, 2, 2 > inverse(const Mat< T, 2, 2 > &m)source#
Mat< T, 3, 3 > inverse(const Mat< T, 3, 3 > &m)source#
Mat< T, 4, 4 > inverse(const Mat< T, 4, 4 > &m)source#

The inverse of a 2×2 matrix, in closed form.

Template parameters
T

the element type; floating-point, since the inverse divides.

Parameters
m

the matrix; must be non-singular.

Returns

m⁻¹.

Parameters
std::domain_error

when m is singular (zero determinant).

Complexity

O(1).

Allocation

none.

fn T distance(const Vec< T, N > &a, const Vec< T, N > &b) source#

The Euclidean distance between two points — norm(a - b).

Template parameters
T

the element type; floating-point, since the result is a root.

N

the dimension.

Parameters
a, b

the points.

Returns

‖a − b‖.

Complexity

O(N).

Allocation

none.

fn T distance_squared(const Vec< T, N > &a, const Vec< T, N > &b) source#

The squared distance between two points — squared_norm(a - b).

Prefer it to distance when only comparing distances: it skips the square root.

Template parameters
T

the element type.

N

the dimension.

Parameters
a, b

the points.

Returns

‖a − b‖².

Complexity

O(N).

Allocation

none.

fn Vec< T, N > reflect(const Vec< T, N > &incident, const Vec< T, N > &normal) source#

Reflect an incident vector about a surface normal — I − 2 (N·I) N, the GLSL reflect.

normal is assumed unit length, as GLSL requires.

Template parameters
T

the element type.

N

the dimension.

Parameters
incident

the incoming vector.

normal

the unit surface normal.

Returns

the reflected vector.

Complexity

O(N).

Allocation

none.

fn Vec< T, N > refract(const Vec< T, N > &incident, const Vec< T, N > &normal, T eta) source#

Refract an incident vector through a surface — the GLSL refract.

incident and normal are assumed unit length. On total internal reflection (a negative radicand) the result is the zero vector, exactly as GLSL specifies.

Template parameters
T

the element type; floating-point.

N

the dimension.

Parameters
incident

the unit incoming vector.

normal

the unit surface normal.

eta

the ratio of indices of refraction (source over destination).

Returns

the refracted vector, or the zero vector under total internal reflection.

Complexity

O(N).

Allocation

none.

fn Vec< T, N > faceforward(const Vec< T, N > &n, const Vec< T, N > &incident, const Vec< T, N > &reference) source#

Orient a normal to face a viewer — the GLSL faceforward: return n when reference points against the incident direction (dot(reference, incident) < 0), -n otherwise.

Used to keep a surface normal on the camera's side.

Template parameters
T

the element type.

N

the dimension.

Parameters
n

the normal to orient.

incident

the incident vector.

reference

the reference normal the result is oriented against.

Returns

n or -n.

Complexity

O(N).

Allocation

none.

fn Fixed< T, Dims... > abs(const Fixed< T, Dims... > &x) source#

The absolute value of every element.

Template parameters
T

the element type; a real number, so < 0 is meaningful.

Dims

the extents.

Parameters
x

the array.

Returns

|x| elementwise.

Complexity

O(size).

Allocation

none.

fn Fixed< T, Dims... > sign(const Fixed< T, Dims... > &x) source#

The sign of every element: -1, 0, or +1.

Template parameters
T

the element type; a real number.

Dims

the extents.

Parameters
x

the array.

Returns

the elementwise sign.

Complexity

O(size).

Allocation

none.

fn min · 2 overloads
Fixed< T, Dims... > min(const Fixed< T, Dims... > &a, const Fixed< T, Dims... > &b)source#
Fixed< T, Dims... > min(const Fixed< T, Dims... > &x, T s)source#

The smaller of each corresponding pair of elements.

Template parameters
T

the element type; a real number.

Dims

the extents.

Parameters
a, b

the arrays.

Returns

min(aᵢ, bᵢ) elementwise.

Complexity

O(size).

Allocation

none.

fn max · 2 overloads
Fixed< T, Dims... > max(const Fixed< T, Dims... > &a, const Fixed< T, Dims... > &b)source#
Fixed< T, Dims... > max(const Fixed< T, Dims... > &x, T s)source#

The larger of each corresponding pair of elements.

Template parameters
T

the element type; a real number.

Dims

the extents.

Parameters
a, b

the arrays.

Returns

max(aᵢ, bᵢ) elementwise.

Complexity

O(size).

Allocation

none.

fn clamp · 2 overloads
Fixed< T, Dims... > clamp(const Fixed< T, Dims... > &x, T lo, T hi)source#
Fixed< T, Dims... > clamp(const Fixed< T, Dims... > &x, const Fixed< T, Dims... > &lo, const Fixed< T, Dims... > &hi)source#

Constrain every element to [lo, hi] — the GLSL clamp with scalar bounds, the common case of pinning a colour to [0, 1].

Template parameters
T

the element type; a real number.

Dims

the extents.

Parameters
x

the array.

lo

the lower bound.

hi

the upper bound.

Returns

min(max(xᵢ, lo), hi) elementwise.

Complexity

O(size).

Allocation

none.

fn mix · 2 overloads
Fixed< T, Dims... > mix(const Fixed< T, Dims... > &a, const Fixed< T, Dims... > &b, T t)source#
Fixed< T, Dims... > mix(const Fixed< T, Dims... > &a, const Fixed< T, Dims... > &b, const Fixed< T, Dims... > &t)source#

Linear interpolation — the GLSL mix: a (1 − t) + b t, with a scalar blend t (0 gives a, 1 gives b).

Composed from the operators, so it inherits their vectorization.

Template parameters
T

the element type; floating-point.

Dims

the extents.

Parameters
a, b

the endpoints.

t

the blend factor.

Returns

the interpolated array.

Complexity

O(size).

Allocation

none.

fn Fixed< T, Dims... > step(T edge, const Fixed< T, Dims... > &x) source#

A step at edge — the GLSL step: 0 where an element is below edge, 1 at or above.

Template parameters
T

the element type; a real number.

Dims

the extents.

Parameters
edge

the threshold.

x

the array.

Returns

xᵢ < edge ? 0 : 1 elementwise.

Complexity

O(size).

Allocation

none.

fn Fixed< T, Dims... > smoothstep(T edge0, T edge1, const Fixed< T, Dims... > &x) source#

A smooth Hermite transition from 0 to 1 across [edge0, edge1] — the GLSL smoothstep, with everything below edge0 giving 0 and everything above edge1 giving 1.

Template parameters
T

the element type; floating-point.

Dims

the extents.

Parameters
edge0

the lower edge.

edge1

the upper edge.

x

the array.

Returns

the smoothstepped array.

Complexity

O(size).

Allocation

none.

fn Mat< T, R, C > matrix_comp_mult(const Mat< T, R, C > &a, const Mat< T, R, C > &b) source#

The elementwise (Hadamard) product — the GLSL matrixCompMult.

Named apart from operator* precisely because * is the matrix product; this multiplies corresponding entries.

Template parameters
T

the element type.

R

the rows.

C

the columns.

Parameters
a, b

the matrices.

Returns

the elementwise product.

Complexity

O(R·C).

Allocation

none.

fn Mat< T, R, C > outer_product(const Vec< T, R > &c, const Vec< T, C > &r) source#

The outer product of a column and a row — the GLSL outerProduct: an R×C matrix whose (i, j) entry is c[i] · r[j].

A rank-one update, the workhorse of a covariance accumulation.

Template parameters
T

the element type.

R

the length of c (the rows).

C

the length of r (the columns).

Parameters
c

the column vector.

r

the row vector.

Returns

the R×C outer product.

Complexity

O(R·C).

Allocation

none.

fn Mat< T, N, N > inverse_transpose(const Mat< T, N, N > &m) source#

The inverse transpose of a matrix — transpose(inverse(m)), the GLSL inverseTranspose.

This is the matrix that carries normals correctly under a non-uniform transform, so lighting stays right.

Template parameters
T

the element type; floating-point.

N

the dimension.

Parameters
m

the matrix; must be non-singular.

Returns

(m⁻¹)ᵀ.

Parameters
std::domain_error

when m is singular (via inverse).

Complexity

O(1) at the fixed sizes.

Allocation

none.

fn Vec< T, C > row(const Mat< T, R, C > &m, Ix i) source#

Extract one row of a matrix as a vector.

Template parameters
T

the element type.

R

the rows.

C

the columns.

Ix

the index type: an integer, or a scoped enum class naming the row.

Parameters
m

the matrix.

i

the row, 0 <= i < R.

Returns

the C-vector of that row.

Complexity

O(C).

Allocation

none.

fn Vec< T, R > column(const Mat< T, R, C > &m, Ix j) source#

Extract one column of a matrix as a vector — a basis vector of the transform.

This is the axis a scoped enum was made to name: column(view, Axis::Right).

Template parameters
T

the element type.

R

the rows.

C

the columns.

Ix

the index type: an integer, or a scoped enum class naming the column.

Parameters
m

the matrix.

j

the column, 0 <= j < C.

Returns

the R-vector of that column.

Complexity

O(R).

Allocation

none.

fn std::string to_string(const Fixed< T, Dims... > &v) source#

Render v the way an NDArray renders — numpy-style nested brackets, each element through the SHARED scalar formatter (so i8/u8 elements print as NUMBERS, f32/f64 plainly, and a complex as a+bj).

A vector is [a, b, c]; a matrix is [[…], […]] in reading (row, column) order — regardless of the column-major storage. This is what io.print/io.str/str() show.

Parameters
v

the value to format.

Returns

the bracketed text.

Complexity

O(Fixed::size).

Allocation

allocates the result string and a formatting stream per element.

fn std::ostream & operator<<(std::ostream &os, const Fixed< T, Dims... > &v) source#

Stream v (the nested-bracket to_string form), so a Fixed is directly Streamable — a cheatah io.print(v) / io.str(v) finds this by ADL, exactly as it does for an NDArray or a primitive.

Parameters
os

the stream.

v

the value.

Returns

os.

Complexity

O(Fixed::size).

Allocation

allocates the intermediate string and a formatting stream per element.

Constants & variables

var std::size_t extent_product source#

The product of a pack of extents — a fixed array's element count, 1 for an empty pack.

Template parameters
Dims

the extents.

Types

type Fixed< T, N > Vec source#

A fixed-extent vector of N elements.

Template parameters
T

the element type.

N

the length.

type Fixed< T, R, C > Mat source#

A fixed-extent matrix of R rows and C columns — column-major storage, mathematical (row, col) indexing (see the file doc).

Template parameters
T

the element type.

R

the rows.

C

the columns.

type Vec< float, 2 > vec2f source#

A 2-D vector of float.

type Vec< float, 3 > vec3f source#

A 3-D vector of float — a direction, a position, a colour.

type Vec< float, 4 > vec4f source#

A 4-D vector of float — a homogeneous point, an RGBA colour.

type Vec< double, 2 > vec2d source#

A 2-D vector of double.

type Vec< double, 3 > vec3d source#

A 3-D vector of double.

type Vec< double, 4 > vec4d source#

A 4-D vector of double.

type Mat< float, 2, 2 > mat2f source#

A 2×2 matrix of float.

type Mat< float, 3, 3 > mat3f source#

A 3×3 matrix of float — a rotation, or a normal matrix.

type Mat< float, 4, 4 > mat4f source#

A 4×4 matrix of float — a transform; exactly the 64 bytes of a push constant.

type Mat< double, 2, 2 > mat2d source#

A 2×2 matrix of double.

type Mat< double, 3, 3 > mat3d source#

A 3×3 matrix of double.

type Mat< double, 4, 4 > mat4d source#

A 4×4 matrix of double.