linalg benchmarks
The generated measurement tables for the linalg module — every case, the harness that measured it, and how to reproduce it.
linalg goes head-to-head with NumPy, whose array ops dispatch to BLAS/LAPACK — hand-tuned, vectorized, often multi-threaded Fortran. numpy_compare.py feeds the same fixed-seed, well-conditioned matrix to both sides, runs the same op many times with the result consumed, and flags any disagreement in the answers.
Read the stamp before the numbers: NumPy's speed depends on which BLAS it links, so the stamp records the resolved library, not the NumPy version. Here that is libblas.so.3 → the reference implementation, not OpenBLAS or MKL — a tuned BLAS would narrow the large-n rows. The small-n wins come from paying no per-call Python and dispatch overhead, which no choice of BLAS changes. The Eigen comparison is a separate Google Benchmark harness (eigen_compare_bench.cpp) with both sides compiled C++ on one thread; the two tables never share a column. Methodology and the reference machine: docs/performance.md.
vs NumPy
Measured by scripts/numpy_compare.py; reproduce with python3 scripts/numpy_compare.py --suite linalg --md docs/bench/linalg-vs-numpy.md.
op | operand dimensions | cheatah | NumPy | winner | band |
|---|---|---|---|---|---|
| 4 | 0.09 | 0.79 | cheatah 8.3x | 7.20-8.48 |
| 16 | 0.54 | 2.24 | cheatah 4.1x | 3.56-4.29 |
| 32 | 3.22 | 13.43 | cheatah 4.2x | 3.94-4.60 |
| 64 | 25.60 | 116.37 | cheatah 4.5x | 4.30-4.71 |
| 96 | 90.65 | 318.41 | cheatah 3.5x | 3.38-3.63 |
| 4 | 0.22 | 2.45 | cheatah 11.0x | 10.03-11.49 |
| 16 | 1.06 | 4.34 | cheatah 4.1x | 3.81-4.35 |
| 32 | 3.69 | 10.35 | cheatah 2.8x | 2.55-2.85 |
| 64 | 18.02 | 50.82 | cheatah 2.8x | 2.72-3.10 |
| 4 | 0.11 | 2.05 | cheatah 22.6x | 15.44-28.51 |
| 16 | 0.81 | 3.76 | cheatah 4.7x | 3.94-5.39 |
| 32 | 2.80 | 9.58 | cheatah 3.4x | 3.18-3.53 |
| 64 | 14.17 | 47.05 | cheatah 3.4x | 3.22-3.51 |
| 4 | 0.26 | 2.18 | cheatah 8.2x | 7.66-9.08 |
| 16 | 1.56 | 6.35 | cheatah 4.1x | 3.18-4.40 |
| 32 | 6.97 | 25.27 | cheatah 3.6x | 3.51-3.65 |
| 64 | 46.20 | 142.12 | cheatah 3.1x | 2.95-3.51 |
| 2 | 0.22 | 2.01 | cheatah 9.2x | 8.95-9.74 |
| 3 | 0.48 | 2.35 | cheatah 4.9x | 4.67-5.01 |
| 4 | 0.62 | 2.66 | cheatah 4.3x | 3.88-4.37 |
| 6 | 1.43 | 3.49 | cheatah 2.4x | 2.34-2.48 |
| 8 | 1.91 | 4.20 | cheatah 2.2x | 2.10-2.23 |
| 16 | 6.73 | 9.20 | cheatah 1.4x | 1.32-1.39 |
| 32 | 26.01 | 25.98 | cheatah 1.0x | 0.95-1.04 |
| 64 | 106.80 | 149.84 | cheatah 1.4x | 1.26-1.43 |
| 64 | 0.02 | 0.62 | cheatah 34.4x | 28.33-46.46 |
| 1024 | 0.09 | 1.08 | cheatah 11.6x | 7.28-17.05 |
| 16384 | 2.59 | 8.16 | cheatah 3.1x | 2.77-3.54 |
| 64 | 0.19 | 0.88 | cheatah 4.7x | 3.07-4.97 |
| 1024 | 0.96 | 1.77 | cheatah 1.8x | 1.52-2.09 |
| 16384 | 14.58 | 14.66 | NumPy 1.0x | 0.95-1.04 |
| 64 | 0.21 | 1.05 | cheatah 5.0x | 3.83-5.62 |
| 1024 | 1.00 | 4.03 | cheatah 4.0x | 2.52-4.76 |
| 16384 | 13.52 | 48.89 | cheatah 3.6x | 3.40-3.92 |
| 64 | 0.22 | 1.11 | cheatah 5.0x | 4.08-6.13 |
| 1024 | 0.95 | 5.52 | cheatah 5.9x | 4.23-6.25 |
| 16384 | 13.69 | 89.32 | cheatah 6.5x | 5.68-6.84 |
| 64 | 0.11 | 0.62 | cheatah 5.5x | 4.58-6.05 |
| 16384 | 2.48 | 2.88 | cheatah 1.2x | 0.78-1.23 |
| 8 | 0.31 | 2.63 | cheatah 8.4x | 7.20-9.79 |
| 32 | 3.77 | 7.35 | cheatah 2.0x | 1.68-2.15 |
| 64 | 16.81 | 25.99 | cheatah 1.5x | 1.35-1.78 |
| 8 | 1.07 | 8.88 | cheatah 8.3x | 7.53-8.64 |
| 32 | 14.79 | 28.49 | cheatah 2.0x | 1.72-2.16 |
| 64 | 102.62 | 137.47 | cheatah 1.3x | 1.16-1.36 |
| 8 | 3.28 | 6.15 | cheatah 1.9x | 1.70-1.93 |
| 32 | 49.70 | 42.83 | NumPy 1.2x | 0.82-0.96 |
| 64 | 229.46 | 208.75 | NumPy 1.1x | 0.88-0.97 |
| 8 | 3.93 | 10.53 | cheatah 2.7x | 2.51-2.85 |
| 32 | 70.11 | 115.62 | cheatah 1.6x | 1.50-1.65 |
| 64 | 433.35 | 739.34 | cheatah 1.7x | 1.50-1.76 |
| 8 | 4.63 | 18.86 | cheatah 4.1x | 3.37-4.19 |
| 32 | 100.36 | 136.64 | cheatah 1.4x | 1.31-1.47 |
| 64 | 633.48 | 800.84 | cheatah 1.3x | 1.23-1.31 |
| 8 | 3.23 | 10.21 | cheatah 3.2x | 3.07-3.31 |
| 32 | 49.56 | 45.71 | NumPy 1.1x | 0.89-0.97 |
| 64 | 217.89 | 195.27 | NumPy 1.1x | 0.89-0.95 |
| 8 | 2.98 | 11.98 | cheatah 4.0x | 3.86-4.16 |
| 32 | 46.95 | 51.17 | cheatah 1.1x | 0.98-1.12 |
| 64 | 216.33 | 194.85 | NumPy 1.1x | 0.88-0.93 |
| 8 | 0.20 | 3.15 | cheatah 15.2x | 11.78-17.67 |
| 32 | 2.93 | 10.48 | cheatah 3.6x | 3.45-3.71 |
| 64 | 14.03 | 48.95 | cheatah 3.5x | 2.20-3.56 |
| 8 | 2.48 | 6.52 | cheatah 2.6x | 2.32-4.14 |
| 32 | 37.96 | 64.51 | cheatah 1.7x | 1.62-1.72 |
| 64 | 238.38 | 401.26 | cheatah 1.7x | 1.61-1.72 |
| 8 | 7.88 | 14.11 | cheatah 1.8x | 1.73-1.89 |
| 8 | 0.91 | 2.49 | cheatah 2.7x | 2.60-2.79 |
| 32 | 12.27 | 28.00 | cheatah 2.3x | 1.96-2.38 |
| 64 | 97.93 | 214.16 | cheatah 2.2x | 2.12-2.40 |
| 32 | 0.01 | 1.20 | cheatah 168.1x | 120.84-178.23 |
| 256 | 0.08 | 1.47 | cheatah 20.2x | 17.90-22.16 |
| 32 | 0.09 | 1.55 | cheatah 17.8x | 10.01-24.71 |
| 256 | 7.38 | 32.35 | cheatah 4.4x | 3.92-4.52 |
| 64 | 0.34 | 4.05 | cheatah 11.9x | 7.13-13.10 |
| 256 | 9.78 | 48.77 | cheatah 5.0x | 4.15-5.25 |
| 8 | 1.71 | 13.12 | cheatah 7.8x | 7.24-8.74 |
| 16 | 20.25 | 72.57 | cheatah 3.6x | 3.20-4.17 |
| 32 | 347.09 | 950.98 | cheatah 2.7x | 2.49-2.81 |
vs Eigen
Measured by tests/benchmarks/eigen_compare_bench.cpp; reproduce with scripts/bench_run.sh publish linalg-vs-eigen.
case | cheatah | spread | vs | rival | spread | ratio | verdict |
|---|---|---|---|---|---|---|---|
BM_dot/16384 | 1.96 µs | ±249.61 ns IQR | eigen | 2.58 µs | ±71.26 ns IQR | 1.32x | faster |
BM_dot/64 | 11.66 ns | ±1.57 ns IQR | eigen | 9.08 ns | ±0.37 ns IQR | 0.78x | slower |
BM_inv/32 | 6.82 µs | ±1.34 µs IQR | eigen | 11.86 µs | ±133.38 ns IQR | 1.74x | faster |
BM_inv/64 | 40.48 µs | ±8.56 µs IQR | eigen | 67.73 µs | ±1.32 µs IQR | 1.67x | faster |
BM_matmul/32 | 3.23 µs | ±1.23 µs IQR | eigen | 3.90 µs | ±115.58 ns IQR | 1.21x | faster |
BM_matmul/96 | 58.11 µs | ±32.71 µs IQR | eigen | 94.56 µs | ±2.47 µs IQR | 1.63x | faster |
BM_solve/32 | 3.59 µs | ±175.70 ns IQR | eigen | 4.49 µs | ±71.85 ns IQR | 1.25x | faster |
BM_solve/64 | 16.73 µs | ±1.91 µs IQR | eigen | 20.77 µs | ±525.30 ns IQR | 1.24x | faster |
Tally (a difference counts only above 1.15x AND 0.25 ns) — vs eigen: 7 faster / 0 parity / 1 slower.
Loss vs eigen:
BM_dot/64— cheatah 11.66 ns vs 9.08 ns (1.28x slower)
What the two tables say
vs NumPy/LAPACK cheatah wins across nearly the whole library — the products, the LU family, the full SVD, the symmetric eigensolver,
outer,qr,kronandnorm. The rows still behind are the SVD-threshold queries —svdvalsandcondat 32 and 64, andmatrix_rankat 64 — where LAPACK's bidiagonal solver edges it, and one elementwise tie.vs Eigen 3.4 on one core cheatah leads
inv,matmul,solveand the largedot, and loses the 64-elementdot, where Eigen's kernel is the tighter one. Those four routines are the whole published Eigen suite, and its tally is its own — the two tables never divide into each other.
