linalg
linalg — with its stamp (commit, host, harness) and the commands that reproduce it — is on the linalg benchmarks page.NumPy-style linear algebra on ndarray, with SIMD-accelerated contiguous kernels.
For the small-and-hot regime — a 3-D direction, a 4×4 transform built and consumed millions of times a second — reach for the sibling fixarray module: the same mathematics with the shape moved into the type (vec3f/mat4f, allocation-free, and never slower than GLM anywhere the two overlap). linalg is the home of the heavy, shape-generic numerics below.
The routines mirror numpy.linalg and operate on ndarray::NDArray (2-D = matrix, 1-D = vector). The general eigensolvers eig/eigvals return a complex spectrum (CNDArray) — a real matrix can have complex conjugate eigenvalue pairs, e.g. a rotation has ±i — while the Hermitian solvers eigh/eigvalsh return a guaranteed-real spectrum, the same split as numpy.
Usage
import ndarray
import linalg # links ndarray for you
let A = ndarray.reshape(ndarray.array([4.0, 1.0, 1.0, 3.0]), [2, 2])
let b = ndarray.array([1.0, 2.0])
let x = linalg.solve(A, b) # A·x = b
let d = linalg.det(A)Functions
Every routine below returns a fresh result. Each also has an out-parameter overload (solve(out, A, b), svd(u, s, vh, A), …) in routines.hpp that writes into caller-owned storage: the products and reductions stay off the heap on contiguous operands, and the factorizations keep their private scratch.
Products
dot/vdot/inner— vector dot product (Σ aᵢbᵢ).outer— outer product of two vectors → matrix.matmul— matrix multiply.matrix_power— integer matrix power Aⁿ (negative viainv).kron— Kronecker (block) product.
Decompositions
cholesky— lower-triangular L with A = L·Lᵀ (SPD only).qr— reduced QR via Householder reflections.svd— singular value decomposition (Golub–Reinsch: bidiagonalization + implicit QR).svdvals— singular values only (the SVD fast path, without forming U/Vᵀ).
Eigenvalues
eig/eigvals— general square matrix (complex spectrum + eigenvectors; Hessenberg + shifted QR for the values, inverse iteration for the vectors).eigh/eigvalsh— symmetric or complex Hermitian matrix (real spectrum; real eigenvectors for a real symmetric matrix, complex eigenvectors for a Hermitian one; Householder tridiagonalization + QL, via a real 2n embedding for Hermitian input).
All eigenvalue routines return the spectrum descending — numpy's eigvalsh returns ascending — with eigenvector columns reordered to match.
import io
import ndarray
import linalg
# A real rotation matrix has complex eigenvalues ±i:
let r = ndarray.reshape(ndarray.array([0.0, -1.0, 1.0, 0.0]), [2, 2])
io.print(ndarray.to_string(linalg.eigvals(r))) # [0+1j, 0-1j]Complex inner-product spaces
Vectors and matrices can be complex (ndarray.complex(re, im)), so the routines work over complex inner-product spaces — Hermitian operators, complex wavefunctions:
dot— bilinear product Σ aᵢbᵢ (complex, no conjugation; matches numpy).vdot— conjugate-linear Hermitian inner product ⟨a, b⟩ = Σ conj(aᵢ)·bᵢ (conjugates the first argument;vdot(a, a)is the real ‖a‖²).matmul— complex matrix multiply.conj_transpose— conjugate transpose (Hermitian adjoint) Aᴴ.
import io
import ndarray
import linalg
# A Hermitian operator H = [[2, 1+i], [1-i, 3]] — real eigenvalues 4, 1:
let re = ndarray.array([2.0, 1.0, 1.0, 3.0])
let im = ndarray.array([0.0, 1.0, -1.0, 0.0])
let H = ndarray.reshape(ndarray.complex(re, im), [2, 2])
io.print(ndarray.to_string(linalg.eigvalsh(H))) # [4, 1]
# Hermitian inner product ⟨a, a⟩ = ‖a‖² is real:
let a = ndarray.complex(ndarray.array([1.0, 3.0]), ndarray.array([2.0, -1.0]))
io.print(linalg.vdot(a, a)) # 15+0jNorms & numbers
norm— L2 (vector) / Frobenius (matrix).cond— 2-norm condition number.det/slogdet— determinant (LU);slogdetis overflow-safe.matrix_rank— numerical rank from SVD thresholding.trace— sum of the main diagonal.
Solving & inverses
solve— solve A·x = b via LU with partial pivoting.lstsq— least-squares solution min‖A·x − b‖.inv— matrix inverse (LU).pinv— Moore–Penrose pseudo-inverse (SVD, any shape).
SIMD
SIMD here is pure compiler auto-vectorization (no intrinsics): contiguous, unit-stride loops compiled at -O3 -march=native. These functions only report the build's capability:
simd_features— instruction sets this build targets (e.g.AVX2;FMA,NEON,scalar).simd_lane_doubles— widest SIMD lane width indoubles.
A build with no SIMD returns identical results, just slower — SIMD is never a correctness dependency. The model, and its compile-time-dispatch limitation, is in simd.hpp.
Performance
linalg goes head-to-head with NumPy, whose array ops dispatch to BLAS/LAPACK — hand-tuned, vectorized, often multi-threaded Fortran — and, in a separate single-core harness, with Eigen 3.4. Each function's Performance row on this page points at the generated tables; the full size-dependence, and the BLAS caveat that governs how to read it, is on the linalg benchmarks page.
cheatah does all of this on one core, by design — no hidden threads, nothing to tune. NumPy's and Eigen's remaining edges are the large blocked problems where threaded BLAS-3 kernels spread the work; the Performance guide has the rationale.
Per-function docs (parameters, complexity, heap behavior) are in routines.hpp. Tested in ../tests/linalg_routines_test.cpp and ../tests/linalg_smoke_test.cpp; ASan + Valgrind clean via the QA gate (security/run-valgrind.sh).
Functions
matmul · 2 overloads
Matrix multiply — the allocating front.
Both operands are Array<T> (so host⊗device / element mixes fail to deduce and are compile errors); requires both to be 2-D with matching inner dimensions — or both 3-D for the BATCHED product [B,M,K] @ [B,K,N] → [B,M,N] (equal batch counts, strict: no broadcast batching). Allocates the result via Array<T>::uninitialized and fills it through the out-parameter kernel — the host SIMD path, or a device shader when Array is a device container.
T | the element type ( |
Array | the container template. |
a | m×k matrix, or a B×m×k batch of matrices. |
b | k×p matrix, or a B×k×p batch. |
m×p product (or the B×m×p batch), an Array<T> of the same container and element.
O(n³) (× B for a batch).
allocates only the result; operands read in place (a strided host view packs once).
deliberately single-threaded (the fastest-per-core contract); parallelize across independent products in the caller.
LinalgCompileRun.MatmulStdlibE2E.Linalg SystemApps.LinearSolveThe flattened length of a vector-shaped operand — 1-D, or 2-D with a size-1 row/column (throws otherwise).
Reads only host-resident shape metadata, so it is valid for ANY located container, device arrays included; the shared validation step of every vector front below.
A | the (located) container type. |
a | the operand whose vector length is wanted. |
the element count of the flattened vector.
O(1).
none.
LinalgRoutines.ProductsAndTracedot · 2 overloads
Dot product: 1-D inner product (vectors flattened) — the bilinear Σ aᵢbᵢ.
Flattens each operand to a vector (1-D, or 2-D with a size-1 row/column) and throws if either is not vector-shaped or the lengths differ. Both operands are Array<T> (the deduction firewall).
T | the element type. |
Array | the container template. |
a, b | same-length vectors. |
Σ aᵢbᵢ as the scalar T.
O(n).
none for contiguous operands (read in place); a non-contiguous view packs once O(n).
LinalgCompileRun.Dot LinalgCompileRun.ComplexDotStdlibE2E.Linalg StdlibE2E.LinalgComplexvdot · 2 overloads
Vector dot product.
For a REAL element this is the bilinear Σ aᵢbᵢ (identical to dot and inner); for a complex element it is the conjugate-linear Hermitian inner product ⟨a, b⟩ = Σ conj(aᵢ)·bᵢ (numpy's vdot, conjugating the first argument); vdot(a, a) is ‖a‖².
T | the element type. |
Array | the container template. |
a, b | same-length vectors. |
Σ aᵢbᵢ (real) or Σ conj(aᵢ)·bᵢ (complex), as the scalar T.
O(n).
none for contiguous operands; a non-contiguous view packs once O(n).
LinalgCompileRun.Vdot LinalgCompileRun.ComplexVdotStdlibE2E.Linalg StdlibE2E.LinalgComplexinner · 2 overloads
Inner product of two vectors — the bilinear Σ aᵢbᵢ (numpy's inner; same as dot).
T | the element type. |
Array | the container template. |
a, b | same-length vectors. |
Σ aᵢbᵢ as the scalar T.
O(n).
none for contiguous operands; a non-contiguous view packs once O(n).
LinalgRoutines.VdotInnerOuterKronLinalgCompileRun.InnerStdlibE2E.Linalgtrace · 2 overloads
Trace: the sum of the matrix diagonal, as the scalar T.
Requires a 2-D matrix (throws otherwise); rectangular matrices sum min(r, c) diagonal entries.
T | the element type. |
Array | the container template. |
a | a 2-D matrix. |
Σ aᵢᵢ as the scalar T.
O(min(r, c)).
none (strided diagonal read straight from the buffer).
LinalgRoutines.ProductsAndTraceLinalgCompileRun.TraceStdlibE2E.Linalg SystemApps.LinearSolveouter · 2 overloads
Outer product of two vectors.
Flattens both operands to vectors and forms the full rank-1 matrix; any pair of vector lengths is accepted (no matching constraint). Allocates the result via Array<T>::uninitialized and fills it through the out-parameter kernel — the host SIMD path, or a device shader when Array is a device container (selected by concept).
a | length-n vector. |
b | length-m vector. |
n×m matrix aᵢbⱼ.
O(n·m).
allocates only the n×m result; operands read in place when contiguous.
LinalgRoutines.VdotInnerOuterKronLinalgCompileRun.OuterStdlibE2E.Linalgconj_transpose · 2 overloads
Conjugate transpose (Hermitian adjoint) Aᴴ: transpose, then conjugate every entry (a plain transpose for a real element — the conjugation is compiled out).
A matrix is Hermitian iff conj_transpose(A) == A.
a | a 2-D matrix. |
the c×r adjoint of an r×c input; throws on non-2-D input.
O(r·c).
allocates only the c×r result; a non-contiguous operand is packed once into scratch.
LinalgRoutines.ComplexProductsLinalgCompileRun.ConjTransposeStdlibE2E.LinalgComplexkron · 2 overloads
Kronecker product.
Requires both operands to be 2-D (throws otherwise) and replaces each entry of a with that scalar times the whole of b, giving the (m·p)×(k·q) block matrix; no dimension matching is needed.
a | m×k matrix. |
b | p×q matrix. |
(m·p)×(k·q) block product.
O(n⁴) in the output area.
allocates only the (m·p)×(k·q) result; a non-contiguous operand packs once into scratch.
LinalgRoutines.VdotInnerOuterKronLinalgCompileRun.KronStdlibE2E.Linalgdot< double, ndarray::basic_ndarray > · 2 overloads
dot< Cplx, ndarray::basic_ndarray > · 2 overloads
vdot< double, ndarray::basic_ndarray > · 2 overloads
vdot< Cplx, ndarray::basic_ndarray > · 2 overloads
inner< double, ndarray::basic_ndarray > · 2 overloads
outer< double, ndarray::basic_ndarray > · 2 overloads
matmul< double, ndarray::basic_ndarray > · 2 overloads
matmul< Cplx, ndarray::basic_ndarray > · 2 overloads
conj_transpose< Cplx, ndarray::basic_ndarray > · 2 overloads
conj_transpose< double, ndarray::basic_ndarray > · 2 overloads
matrix_power · 2 overloads
Integer matrix power Aⁿ (negative n via inv).
Requires a square matrix (throws otherwise); n == 0 returns the identity, and negative n first inverts a via inv (so it inherits inv's singular-matrix behavior) before raising to |n|.
a | square matrix. |
n | exponent. |
Aⁿ.
O(n³·log|n|) by binary exponentiation.
allocates a new NDArray result; the binary-exponentiation matmul steps allocate their own intermediates.
LinalgRoutines.MatrixPowerLinalgCompileRun.MatrixPowerStdlibE2E.Linalgkron< double, ndarray::basic_ndarray > · 2 overloads
trace< double, ndarray::basic_ndarray > · 2 overloads
norm · 2 overloads
Norm: L2 for vectors, Frobenius for matrices.
Dispatches on rank: 1-D (or lower) inputs get the Euclidean L2 norm, 2-D inputs the Frobenius norm; either way it is the square root of the sum of squared entries. Two-layer over the element and container like every routine (the scalar-out kernel norm(out, a) is the host/device seam; host double is the shipped instantiation).
a | vector or matrix. |
√Σ xᵢ².
O(n) for vectors / O(n²) for matrices.
none for a contiguous operand (summed in place); a non-contiguous view packs once. Returns a double.
LinalgRoutines.NormAndRankLinalgCompileRun.NormStdlibE2E.Linalg SystemApps.LinearSolvenorm< double, ndarray::basic_ndarray > · 2 overloads
solve · 2 overloads
Solve A·x = b via LU with partial pivoting.
Factorizes a once then does forward/back substitution against b; requires a square and b a vector of matching length (throws otherwise). A singular a does not throw but yields a garbage/overflowing solution (pivots are nudged off zero rather than detected).
a | square coefficient matrix. |
b | right-hand-side vector. |
solution x.
O(n³).
allocates a new NDArray result; the LU factorization allocates its own O(n²) scratch.
LinalgRoutines.SolveDetInvLinalgCompileRun.SolveStdlibE2E.Linalg SystemApps.LinearSolvedet · 2 overloads
Determinant via LU with partial pivoting.
Computes the product of the LU pivots times the permutation sign; requires a square matrix (throws otherwise). A singular matrix yields a determinant of (or extremely near) zero rather than an error.
a | square matrix. |
det(A).
O(n³).
allocates scratch O(n²) for the factorization; returns a double.
LinalgRoutines.SolveDetInvLinalgCompileRun.DetStdlibE2E.Linalg SystemApps.LinearSolveslogdet · 2 overloads
Sign and log|det| via LU (overflow-safe determinant).
Sums the logs of the absolute LU pivots (avoiding the over/underflow of a raw product) and tracks the sign from the pivot signs and permutation parity; requires a square matrix (throws otherwise). A singular matrix gives a hugely negative logabsdet rather than −infinity, since a zero pivot is nudged to a tiny value during factorization.
a | square matrix. |
SLogDet.
O(n³).
allocates scratch O(n²) for the factorization (the struct members are plain doubles).
LinalgRoutines.SlogdetAndCondLinalgCompileRun.SlogdetStdlibE2E.Linalginv · 2 overloads
Matrix inverse via LU with partial pivoting.
Factorizes a once and back-solves against each identity column; requires a square matrix (throws otherwise). A singular a does not throw but produces garbage/overflowing entries since zero pivots are nudged rather than detected.
a | square matrix. |
A⁻¹.
O(n³).
allocates a new NDArray result; the LU factorization and the whole-identity back-solve allocate their own O(n²) scratch.
LinalgRoutines.SolveDetInvLinalgCompileRun.InvStdlibE2E.Linalgsolve< double, ndarray::basic_ndarray > · 2 overloads
det< double, ndarray::basic_ndarray > · 2 overloads
slogdet< double, ndarray::basic_ndarray > · 2 overloads
inv< double, ndarray::basic_ndarray > · 2 overloads
lstsq · 2 overloads
Least-squares solution min‖A·x − b‖ (computed as pinv (a)·b).
Forms the Moore–Penrose pseudo-inverse via SVD and multiplies it by b, so it handles over- and under-determined systems and returns the minimum-norm solution for rank-deficient a; b must be conformable for the matmul step.
a | m×n matrix. |
b | right-hand side. |
minimizing x.
iterative O(n³) via SVD.
allocates a new NDArray result; plus the intermediate n×m pseudo-inverse and its SVD scratch.
LinalgRoutines.LstsqLinalgCompileRun.LstsqStdlibE2E.Linalgcholesky · 2 overloads
Cholesky factor of a symmetric positive-definite matrix (throws otherwise).
Requires a square matrix and computes L column by column reading only the lower triangle of a; if any pivot (the diagonal under the square root) is non-positive it throws "matrix is not positive-definite", which also catches non-SPD or non-symmetric input.
a | square SPD matrix. |
lower-triangular L with A = L·Lᵀ.
O(n³).
allocates a new NDArray result; the factor is computed into O(n²) private scratch, then copied in.
LinalgRoutines.CholeskyAndQRLinalgCompileRun.CholeskyStdlibE2E.Linalgqr · 2 overloads
Reduced QR via Householder reflections (requires rows ≥ cols).
Applies successive Householder reflectors to triangularize a, returning the thin/reduced factors; throws "qr requires rows >= cols" for wide matrices. Rank-deficient columns (zero pivot norm) are skipped, leaving the corresponding R entries zero.
a | m×n matrix. |
QR with q (m×n, orthonormal cols) and r (n×n, upper-triangular).
O(m²·n) — n reflectors each applied to the full m×m Q; O(n³) when square.
allocates both members; the factorization works in O(m²) private scratch (a full m×m Q workspace), then copies the reduced factors in.
LinalgRoutines.CholeskyAndQRLinalgCompileRun.QrStdlibE2E.Linalgsvd · 2 overloads
Singular value decomposition (Golub–Reinsch; requires rows ≥ cols).
Reduces a to upper-bidiagonal form by Householder reflections, then diagonalizes it with implicit-shift QR (accumulating U and V), and sorts the singular values descending — the world-standard dense SVD (what LAPACK's dgesvd reduces to). Throws "svd requires rows >= cols" for wide matrices (transpose first); singular values come out non-negative.
a | m×n matrix. |
SVD with u (m×n), s (descending singular values), vh (n×n = Vᵀ).
iterative O(m·n²); O(n³) when square.
allocates all members; the Golub–Reinsch reduction allocates its own O(m·n) workspace.
LinalgRoutines.SvdAndEighLinalgCompileRun.SvdStdlibE2E.Linalgsvdvals · 2 overloads
Singular values only (≈ numpy.linalg.svd(a, compute_uv=False) / svdvals).
Runs the same Golub–Reinsch reduction as svd but takes the values-only fast path — it never accumulates U or V, and skips the U/V Givens rotations in the QR sweep — so it beats the full decomposition by a margin that grows with n (see the linalg benchmark page). Accepts any shape (singular values of a and aᵀ coincide).
a | m×n matrix. |
length-min(m,n) vector of singular values, descending.
iterative O(n³), but a large constant factor below svd.
allocates a new NDArray result; the reduction still allocates its O(m·n) workspace (the values-only path skips the U/V accumulation work and result copies, not the working buffers).
LinalgRoutines.SvdAndEighLinalgCompileRun.SvdvalsStdlibE2E.Linalgpinv · 2 overloads
Moore–Penrose pseudo-inverse via SVD (any shape).
Computes V·diag(1/σ)·Uᵀ from a Golub–Reinsch SVD, transposing wide matrices internally so any shape works; singular values at or below a size-scaled tolerance are dropped (treated as zero) so it stays well-defined for rank-deficient input.
a | m×n matrix. |
n×m pseudo-inverse.
iterative O(n³) via SVD.
allocates a new NDArray result; the SVD and the assembly allocate their own O(m·n) scratch.
LinalgCompileRun.PinvStdlibE2E.Linalgcond · 2 overloads
2-norm condition number σ_max/σ_min (∞ if singular).
Takes the ratio of largest to smallest singular value from a Golub–Reinsch SVD (transposing internally for wide matrices); returns +infinity when the smallest singular value is exactly zero (singular/rank-deficient).
a | matrix. |
condition number.
iterative O(n³) via SVD.
allocates scratch O(n²) for the factorization; returns a double.
LinalgRoutines.SlogdetAndCondLinalgCompileRun.CondStdlibE2E.Linalgmatrix_rank · 2 overloads
Numerical rank from SVD singular-value thresholding.
Counts singular values above a tolerance scaled by the largest singular value and the matrix size (the standard numpy-style threshold); accepts any shape, transposing wide matrices internally.
a | matrix. |
rank.
iterative O(n³) via SVD.
allocates scratch O(n²) for the factorization.
LinalgRoutines.NormAndRankLinalgCompileRun.MatrixRankStdlibE2E.Linalgeigh · 2 overloads
void eigh(Array< ndarray::real_base_t< T > > &values, Array< T > &vectors, const Array< T > &a)source#eigh_result_t< T, Array > eigh(const Array< T > &a)source#Eigen-decomposition of a symmetric matrix (Householder tridiagonalization + QL).
Reduces a to tridiagonal form by Householder reflections, then diagonalizes it with implicit-shift QL, returning real eigenvalues sorted descending with matching eigenvector columns; it reads the full matrix and assumes symmetry rather than checking it, so asymmetric input yields meaningless results. Throws on a non-square matrix (or if the QL iteration fails to converge).
a | square symmetric matrix. |
Eig (real input) or EighC (complex Hermitian input): descending real values and matching eigenvector columns.
iterative O(n³).
allocates both members; the solver allocates its own O(n²) scratch (a complex Hermitian input first embeds into a 2n×2n real matrix).
LinalgCompileRun.Eigh LinalgCompileRun.EighComplexStdlibE2E.Linalg StdlibE2E.LinalgComplexeigvalsh · 2 overloads
Array< ndarray::real_base_t< T > > eigvalsh(const Array< T > &a)source#Eigenvalues of a symmetric matrix, descending (tridiagonal QL).
Same tridiagonalization + QL as eigh but skips the eigenvector accumulation entirely, so it runs well under eigh's cost at large n (see the linalg benchmark page); assumes (does not verify) symmetry and throws on a non-square matrix. ONE two-layer template collapsing the former real and complex (Hermitian) overloads: a real element takes the symmetric path, a complex element the Hermitian path (if constexpr). The spectrum is always REAL, returned as Array<real_base_t<T>>.
T | the element type ( |
Array | the container. |
a | square symmetric (real) / Hermitian (complex) matrix. |
length-n vector of real eigenvalues.
iterative O(n³).
allocates a new result; the solver allocates its own O(n²) scratch (a complex Hermitian input first embeds into a 2n×2n real matrix).
LinalgCompileRun.Eigvalsh LinalgCompileRun.EigvalshComplexStdlibE2E.Linalg StdlibE2E.LinalgComplexeigvalsh< double, ndarray::basic_ndarray > · 2 overloads
eigvalsh< Cplx, ndarray::basic_ndarray > · 2 overloads
eigh< double, ndarray::basic_ndarray > · 2 overloads
eigh< Cplx, ndarray::basic_ndarray > · 2 overloads
eig · 2 overloads
void eig(Array< ndarray::complex_of_t< T > > &values, Array< ndarray::complex_of_t< T > > &vectors, const Array< T > &a)source#EigC< Array< ndarray::complex_of_t< T > > > eig(const Array< T > &a)source#Eigen-decomposition of a general square matrix (complex spectrum and eigenvectors).
For a symmetric a it delegates to eigh (promoted to complex with zero imaginary part); otherwise it uses Hessenberg reduction + shifted QR for the eigenvalues, then inverse iteration for each eigenvector. A real matrix with a complex conjugate pair (e.g. a rotation) yields those complex eigenvalues and eigenvectors rather than throwing. Throws on a non-square matrix or if the QR iteration fails to converge.
a | square matrix. |
EigC with complex values and matching complex eigenvector columns.
iterative O(n³) via Hessenberg + shifted QR for the eigenvalues; the general (non-symmetric) eigenvectors add O(n⁴) — one inverse iteration, each with its own O(n³) complex LU factorization, per eigenvalue (a symmetric a stays O(n³) via eigh, which accumulates the vectors in the QL sweep).
allocates both members; plus O(n²) factorization scratch (on the non-symmetric path, a fresh complex n×n LU per eigenvalue).
LinalgRoutines.GeneralEigLinalgCompileRun.EigStdlibE2E.Linalgeigvals · 2 overloads
Array< ndarray::complex_of_t< T > > eigvals(const Array< T > &a)source#Eigenvalues of a general square matrix (complex), descending.
Routes symmetric input through tridiagonal QL and everything else through Hessenberg + shifted QR, then sorts the result descending (by real part, then by imaginary part). A real matrix with a complex conjugate pair yields those complex eigenvalues rather than throwing. Throws on a non-square matrix or non-convergence of the QR iteration.
a | square matrix. |
length-n complex vector of eigenvalues.
iterative O(n³).
allocates a new CNDArray result; the iteration allocates its own O(n²) scratch.
LinalgRoutines.SvdAndEighLinalgCompileRun.EigvalsStdlibE2E.Linalgmatrix_power< double, ndarray::basic_ndarray > · 2 overloads
lstsq< double, ndarray::basic_ndarray > · 2 overloads
cholesky< double, ndarray::basic_ndarray > · 2 overloads
svdvals< double, ndarray::basic_ndarray > · 2 overloads
pinv< double, ndarray::basic_ndarray > · 2 overloads
cond< double, ndarray::basic_ndarray > · 2 overloads
matrix_rank< double, ndarray::basic_ndarray > · 2 overloads
qr< double, ndarray::basic_ndarray > · 2 overloads
svd< double, ndarray::basic_ndarray > · 2 overloads
eig< double, ndarray::basic_ndarray > · 2 overloads
eigvals< double, ndarray::basic_ndarray > · 2 overloads
Instruction sets this build targets, e.g.
"AVX2;FMA" (x86-64), "NEON" (ARM), or "scalar".
Reflects compile-time target flags (e.g. -march=native), not a runtime CPUID probe.
;-separated feature list.
O(1).
allocates the result string.
LinalgSmoke.SimdFeaturesReportedWidth, in doubles, of the widest SIMD lane this build targets.
1 = scalar, 2 = SSE2/NEON, 4 = AVX, 8 = AVX-512; useful for sizing blocked kernels.
lane width.
O(1).
none.
LinalgSmoke.SimdFeaturesReportedTypes
Whether a reduction conjugates its first operand.
element_t: the scalar an array-like stores — its nested value_type, with any reference/cv-qualification on A stripped first so const NDArray& and NDArray yield the same element type.
Undefined for a type with no value_type (which simply makes the concepts below unsatisfied for it, never a hard error).
location_t: shorthand for the location tag of A (its location_of ::type), with any reference/cv-qualification stripped from A first.
A complex scalar (std::complex<double>) — the element type of CNDArray and the return type of the complex inner products dot / vdot.
A complex array (basic_ndarray<std::complex<double>>) — what the general eigensolvers return, since a real matrix can have complex eigenvalues.
Prints element-wise as a+bj via cheatah::ndarray::to_string.
std::conditional_t< ndarray::is_complex_v< T >, EighC< Array< ndarray::real_base_t< T > >, Array< T > >, Eig< Array< T > > > eigh_result_t
source#
The result type of the unified eigh: for a real element, Eig<Array<T>> (real values + vectors); for a complex element, EighC<Array<real>, Array<T>> (real values, complex vectors).
eigh's return type differs per element, so it is expressed here (not auto) so the header declaration knows it without seeing the definition.
