Skip to content

Performance

The production shape is a sampled decode, which reads a handful of site gridpoints out of a multi-megapoint field. For that shape, decodeJ2kRegion entropy-decodes only the codeblocks the requested points touch and runs window-bounded inverse lifts. No whole-image codec can do this. The values it returns are bit-identical to the full decode, and Two-ring correctness states the contract.

These timings were measured single-threaded (minimum of 5, Node 24, Apple M5 Max; 4 uniformly scattered points per field; recorded 2026-08-12).

fixturesamplesbitsfull decoderegion, 4 pointsratiocodeblocks touched
gdps-tmp-2m2,882,40012266 ms22 ms11.9×49/791
hrdps-continental-tmp-2m3,276,60016692 ms44 ms15.6×49/911
raqdps-pm25-sfc436,67120101 ms20 ms5.0×28/136
rdps-cape-sfc-jasper813,27524295 ms4.4 ms67×25/12,711

The 67× row sits outside the trend of the other rows. rdps-cape-sfc-jasper has the JasPer field’s one-row bitmapped geometry (the JasPer story). An 813275×1 image splits into 12,711 tiny codeblocks, so four points touch a far smaller fraction of them than on any gridded field.

The cost scales sublinearly in points, because nearby points share windows and every point’s coarse-level ancestry converges. On HRDPS-continental (same machine, single thread), 1 point touches 16 codeblocks (14 ms), 4 points 49 (44 ms), 16 points 145 (121 ms), 64 points 366 (297 ms), and 256 points 674 (538 ms). The full packet-structure parse is unavoidable, because packet lengths are only discoverable sequentially. It takes under half a millisecond on every regular fixture.

Region-decode cost tracks codeblocks, not points Three stacked panels. First, the 2540 × 1290 HRDPS continental field with the bench's 4 scattered sample points marked. Next, the codestream's tile buffer: all 911 codeblocks drawn across 5 decomposition levels, with the 49 codeblocks this 4-point decode actually entropy-decoded filled and hatched, clustering around each point's coefficient position in every subband. Below, the touched count as the same scatter grows: 1 point -> 16, 4 points -> 49, 16 points -> 145, 64 points -> 366, 256 points -> 674 of 911 codeblocks; points ×256 costs only ×42 in codeblocks.

Through @azohra/meteo.grib’s worker pool, this is the sampled path end-to-end. On the same machine, a 4-point sampled decode of HRDPS-continental through sampleFieldValuesAsync and a 2-worker pool sustains 25.5 ms per field, against 394 ms per field for full decodes. That is 15.5× per core. A 3,500-field HRDPS lane therefore projects to ~90 s of sampled decode at pool 2, where full decodes would need ~23 minutes. grib/test/production-codec-throughput.test.ts gates the mechanism at ≥6× per core and at bit-exactness against the full decode.

Decoding whole images single-threaded, this decoder is slower than the WASM OpenJPEG build and faster than the asm.js one. The cross-codec bench (grib/tools/bench-j2k-single.ts) measured these timings (minimum of 5, Node 24, Apple Silicon; recorded 2026-08-12). It lives in grib because it needs both codecs, and only that package depends on both.

fixturesamplesbits@azohra/meteo.j2koracleratio
gdps-tmp-2m2,882,40012278 ms98 ms2.84×
hrdps-continental-tmp-2m3,276,60016723 ms282 ms2.57×
raqdps-pm25-sfc436,67120105 ms140 ms0.75×
twelve-fixture total2110 ms883 ms2.39×

This decoder is already faster outright on the 20-bit field. The WASM build clamps samples wider than 16 bits, so its former stand-in on that field was the far slower asm.js artifact. None of the WASM numbers apply to the production shape. The WASM codec decodes whole images only, so a sampled decode under it pays the full-decode column every time.

Profiling puts ~95% of a full decode in Tier-1 (the MQ/EBCOT bit loops; the DWT is ~4%). Tier-1 runs per codeblock, so skipping codeblocks skips the cost, and that is why region decode wins. Tier-1 is also the work src/parallel.ts decomposes. The largest field is 911 independent codeblock tasks, each a pure function over its own byte slice. One full decode can therefore also fan across a worker pool within the field. Neither WASM nor native OpenJPEG can use that dimension at all.

This package uses no Node APIs, so the pool wiring lives in @azohra/meteo.grib/j2k-node. Under its strategy: "codeblock", a full decode of the largest field drops from 736 to 133 ms through an 8-worker pool. See JPEG 2000 and the pool for the pool’s API, sizing, and heap behaviour.

Two benches ship with the packages. One times single-thread full decodes over the corpus (tools/bench.ts). The other is the region bench behind J2K_REGION_BENCH=1 in test/region.test.ts.