cheatah
Module

gpu::dispatch

cheatah-gpu v0.5.1-alpha — Biome Standard 0.6.5-alpha

The backend-agnostic core every GPU compute path needs: turn a problem size into a workgroup launch. Pure C++20 integer arithmetic — header-only, allocation-free, zero dependencies, no platform headers — so it builds, tests, and documents to 100% on a machine with no GPU. It is the CPU-side math both shipped backends consume: the Vulkan surface (see ../vulkan) feeds it straight into vkCmdDispatch, the Metal backend (see ../metal) into MTLComputeCommandEncoder dispatches.

Status: working — workgroup-count arithmetic and the device-limit clamp, pure cheatah with no device required.

import gpu.dispatch as dispatch

# cover 1,000,000 items at local_size_x = 256
let groups = dispatch.group_count_1d(1000000, 256)        # 3907
let safe   = dispatch.clamp_group_count(groups, 65535)    # clamp to maxComputeWorkGroupCount[0]

# a 2-D image dispatch: 1920x1080 pixels at 16x16 threads per group
let g = dispatch.group_count_3d(dispatch.Dim3({.x = 1920, .y = 1080}),
                                dispatch.Dim3({.x = 16,   .y = 16}))   # {120, 68, 1}

Why uint32_t everywhere

Workgroup counts and local sizes are 32-bit unsigned on the hardware: vkCmdDispatch(uint32_t, uint32_t, uint32_t) and VkPhysicalDeviceLimits::maxComputeWorkGroupCount[3] are all uint32_t. We never widen to 64-bit for dimensioning — it doesn't map to the dispatch ABI and wastes shader registers.

API

function

meaning

ceil_div(numerator, denom)

groups to cover numerator items at denom/group (overflow-safe; denom==0 → 0)

group_count_1d(items, local_size)

the groupCountX for a 1-D dispatch

clamp_group_count(want, device_max)

clamp one axis to the device's workgroup-count limit

Dim3{x, y, z}

3-D dispatch extents; unused axes default to 1 (one slice, not zero)

group_count_3d(items, local_size)

the (groupCountX, groupCountY, groupCountZ) for a 3-D dispatch — ceil_div per axis

clamp_group_count(want, device_max) (Dim3 overload)

clamp all three axes to maxComputeWorkGroupCount[3]

Tests

  • Unit (C++ GoogleTest): ../../tests/dispatch_test.cpp — drives 100% line + function coverage.

  • Compile-run (cheatah .purr, one per function): ../../systests/test_dispatch_cr_*.purr — each compiles via purrc against the imported library and must print RESULT: PASS.

  • System (cheatah .purr): ../../systests/test_dispatch.purr, test_dispatch_limits.purr — exercise import gpu.dispatch end-to-end.

Every public function documents @param, @return, @complexity, @alloc, and a @test / @crtest / @systest; the QA gate enforces 100% Javadoc + 100% coverage.

Classes

Functions

fn std::uint32_t ceil_div(std::uint32_t numerator, std::uint32_t denom) #

Number of workgroups needed to cover numerator items at denom items per group (ceiling division).

Computed overflow-safe — without the usual (n + d - 1) / d, which overflows for numerator near UINT32_MAX.

Parameters
numerator

total items to process (e.g. elements in a buffer).

denom

items handled per workgroup — the shader's local size on this axis.

Returns

ceil(numerator / denom); 0 when denom is 0 (an empty/invalid dispatch).

Complexity

O(1).

Host allocation

none.

Unit testDispatch.CeilDiv
Compile-run testsystests/test_dispatch_cr_ceil_div.purr
System testsystests/test_dispatch.purr
fn std::uint32_t group_count_1d(std::uint32_t items, std::uint32_t local_size) #

One-dimensional workgroup count for a 1-D compute dispatch: the groupCountX you pass to vkCmdDispatch so items invocations are covered at local_size threads per group.

Parameters
items

total invocations the shader must cover.

local_size

the shader's local_size_x (threads per workgroup).

Returns

the number of workgroups to dispatch on X; 0 when local_size is 0.

Complexity

O(1).

Host allocation

none.

Unit testDispatch.GroupCount1d
Compile-run testsystests/test_dispatch_cr_group_count_1d.purr
System testsystests/test_dispatch.purr systests/test_dispatch_limits.purr
fn clamp_group_count · 2 overloads
std::uint32_t clamp_group_count(std::uint32_t want, std::uint32_t device_max)#
Dim3 clamp_group_count(Dim3 want, Dim3 device_max)#

Clamp a desired workgroup count for one axis to the device's limit, so a dispatch never exceeds VkPhysicalDeviceLimits::maxComputeWorkGroupCount[axis].

Parameters
want

the workgroup count the problem size asks for.

device_max

the device's maximum workgroup count on this axis.

Returns

want when it fits, otherwise device_max.

Complexity

O(1).

Host allocation

none.

Unit testDispatch.ClampGroupCount
Compile-run testsystests/test_dispatch_cr_clamp_group_count.purr
System testsystests/test_dispatch_limits.purr
fn bool operator==(Dim3 a, Dim3 b) #

Per-axis equality of two extents (all of x, y, and z match).

Parameters
a, b

the extents to compare.

Returns

true when every axis is equal.

Complexity

O(1).

Host allocation

none.

Unit testDispatch.Dim3Equality
Compile-run testsystests/test_dispatch_cr_dim3.purr
fn Dim3 group_count_3d(Dim3 items, Dim3 local_size) #

Three-dimensional workgroup count for a compute dispatch: the (groupCountX, groupCountY, groupCountZ) you pass to vkCmdDispatch so an items volume is covered at local_size threads per group on each axis — ceil_div applied per axis, with its overflow safety and 0-on-zero-denominator semantics.

Parameters
items

total invocations to cover on each axis (e.g. an image's width × height).

local_size

the shader's (local_size_x, local_size_y, local_size_z).

Returns

the per-axis workgroup counts; an axis is 0 when its local_size axis is 0.

Complexity

O(1).

Host allocation

none.

Unit testDispatch.GroupCount3d
Compile-run testsystests/test_dispatch_cr_group_count_3d.purr
System testsystests/test_dispatch.purr