cheatah
Releases

cheatah-space

Every notable change, release by release.

All notable changes to cheatah-space. This project is alpha — expect breaking changes between releases. It is a cheatah standard-library extension and joins the Biome Standard alongside the other cheatah extensions.

Unreleased

space.cdf — reads real NASA CDF files into ndarrays

space.cdf opens a CDF 3.x file and hands its variables back as cheatah ndarrays, GZIP included. That is the whole of this tranche's goal: the flow from a file on disk to a plot.

let f = cdf.open("omni_hro2_1min_20150101_v01.cdf")
let b = cdf.values(f, "F")                  # ndarray[float64], IMF magnitude in nT
let fig = figure.line(figure.new_figure(), hours, b)
plot.save(fig, "out/omni_imf.png")

Verified against bytes NASA produced, not against our own opinion of them. The expected values were decoded by hand from the OMNI byte stream before the reader existed: F opens 6.92 / 5.84 / 5.71 nT and Epoch at 63587289600000.0 ms. And test_alltypes.cdf carries Longitude (GZIP level 9) beside longitude_copy (uncompressed) holding the same values — they now agree value for value, which says our inflate agrees with the encoder NASA's writer used.

Four things learned the hard way, recorded so they are not rediscovered:

Deferred, and refused by name rather than guessed at: CDF 2.x, whole-file compression, multi-file CDFs, VAX floats, sparse records with gaps, RLE/Huffman/adaptive Huffman, column-major N-D variables, the attributes API, the writer, the checksum, signing, and the byte-for-byte differential harness. The three legacy codecs are last on purpose: the specification names them but never documents their bitstreams, and no file in the archive uses any of them.

space.irbem — four external field models, and a vectorised CPU lane

kext=1 Mead & Fairfield (1975), kext=5 Olson-Pfitzer quiet (1977), kext=6 Olson-Pfitzer dynamic (1988) and kext=8 Ostapenko & Maltsev (1997) join T89, each written to its published paper and verified against the vendored Fortran oracle across the four corpus regimes and the four real storm events. Each ships the same shape: a host evaluator templated on the float type, per model validity through status.hpp, a shared Slang function outside every kernel guard so a tracer calls the same physics as a direct evaluation, a guarded kernel, and a registry row that a completeness test checks against the Slang entry points and the CMake shader table.

Mead reaches oracle parity at 2.1e-9 relative RMS (870 grid points x 4 Kp bins x 3 tilts, worst single point 1.2e-7). Establishing that took a black-box provenance experiment rather than a transcription: a free 20-term quadratic fit to the isolated oracle field converges only in SM coordinates at a 4-degree-aberrated position, and refitting on the paper's own y-symmetric basis recovers all 68 published coefficients as 3-4 significant-figure decimals to 1e-9 relative.

Olson-Pfitzer dynamic carries a documented gap, not parity. Its 1988 citation is a conference abstract with no equations and the OP77 parent is a McDonnell Douglas report: neither the functional form nor the coefficients were ever published. What ships is the documented structure -- B = s^3 B_q(s r) + dDst* R(r) with s = (P/P0)^(1/6) -- with every constant from a citable source (CODATA, Chapman-Ferraro/Mead 1964, O'Brien & McPherron 2000, Dessler-Parker-Sckopke). It sits ~50% RMS-relative from IRBEM's kext=6 in the belts, with a measured structure floor of 67.8%. The oracle's tail does follow the published pressure law (fitted s = 1.110/1.260/1.420 /1.475 against a predicted 1.122/1.260/1.414/1.468), but inside 6.5 Re it is an amplitude scaling that no published relation predicts. Parity is unreachable without reading the LGPL source, which this clean room does not do. div B = 0 holds regardless: the stencil residual falls as h^2.

Every model is checked for divergence-free-ness by a second-order stencil whose residual must fall as h^2 -- the one correctness check that needs no oracle and cannot be satisfied by agreeing with a wrong reference.

space.irbem — the CPU batch lane vectorises across points

The Legendre recursion cannot vectorise *within* one field evaluation (loop-carried, ~7 wide), so the batch lane was 84% scalar. batch_soa.hpp makes the point index the SIMD lane and keeps the n/m recursion loops outer: 362 -> 100 ns/point, 3.61x, bit-identical to the scalar lane by memcmp across truncations, epochs, policies and tail lengths. An objdump test counts packed double ops in the strip symbol and fails if the lane ever decays to scalar (a scalar-row control takes it from 390 packed to 0). Bit identity is conditional on -ffp-contract=off, which the repo sets globally and the header now documents for consumers who build outside it.

Three variants measured flat or worse and were reverted rather than shipped: autovectorised plain arrays (2.0x), 512-bit vector rows (0.65x -- double-pumped on AVX2), and 16-point strips (3.3x).

space.cdf — the format layer

The three headers everything above the bytes stands on. No records are parsed yet; this is what makes parsing them safe and testable.

Verified against a real NASA file, not against our own parser's opinion of one. The expected values were decoded by hand from the OMNI hro2_1min byte stream before this code existed: CDR size 312, GDR at 320 = 8 + 312, encoding NETWORK, flags row-major + single-file, zVDRhead 21601, the Epoch zVDR at CDF_EPOCH/maxRec 44639, NumElems at field offset 64 — and eof equal to the file size exactly, the invariant that holds only when no checksum is present.

Two traps found and recorded rather than patched over. mapping.hpp first measured 38% because the real-file tests were skipping: ctest runs the suite from the repo root and scripts/coverage.sh from build/cov, so a bare relative path resolved in one and not the other. A skipped test reads green, so the coverage run had quietly stopped exercising the real syscalls at all — the fake was covered and the production path was not. The corpus path is now searched upward from the working directory. And reaching the real mmap-failure branch needed something that opens and stats plausibly but cannot be mapped: a directory does exactly that.

space.irbem — the TOTAL field runs on the GPU, and storms are verified against the reference

The stated scope limit — "a TotalField batch runs on the CPU today" — is fixed, and the loop the storm corpus was built for is closed against the Fortran oracle.

space.irbem — the frame layer

Work has begun on space.irbem, a from-scratch reimplementation of PRBEM/IRBEM written to the published papers. This first piece is the vocabulary the rest of the module is built from; no field model ships yet.

The IRBEM oracle, and a measured error budget

Repo

v0.1.0-alpha (2026-08-14) — the time module, house-gated

The first release-shaped state of the repo: one working module, space.time (Julian Date, Modified Julian Date, the J2000 epoch offsets, and the NASA CDF_EPOCH bridge), hand-authored as header-only, concept-templated C++20 — scalar and ndarray-vectorized through the same functions — plus the full extension-grade QA harness around it. space.cdf and space.irbem remain roadmaps (design notes in their subdirectory READMEs); nothing of them ships here.

The module

Biome packaging

QA gate grown to the extension bar

scripts/qa_gate.sh (pre-push enforced) now runs: 100% unit coverage (clang source-based, lines + functions over the space headers, README table drift-checked) → 100% Javadoc (scripts/doc_coverage.sh) + compiling doc examples → module-sidecar verification → debug build → the .purr systests (moved from tests/ to systests/) → the biome-install sandbox → GoogleTest unit suite (tests/space_time_test.cpp, new) → ASan + UBSan → Valgrind memcheck (security/run-valgrind.sh) → cppcheck → the private-reference scan (also guarding commit messages via the commit-msg and pre-push hooks).

Removed

Roadmap (unchanged)