Benchmarks
Performance diagnostics and CI correctness coverage.
The current regression suite proves lazy history access, turn-end passivation, isolated subscription backlogs, and native Tool dispatch at full parent concurrency. Previous headline numbers predate this design and are archived in BENCHMARKS.md. They are not release thresholds for the current implementation.
Measure cold readiness, explicit admission, prepared creation, first activation, resumed activation, and history reads separately. Include both the server and Wasmtime worker in memory and CPU totals. Use retained sessions with long histories; an empty-server RSS smoke cannot establish session density.
The existing journal probe reports canonical reopen, disposable checkpoint write, and cached reopen
on the same workload. Run cargo test --release -p brain --test journal_throughput -- --ignored --nocapture.
A checkpoint saves decoding canonical history, but its index and current transcript still grow with
history. Disk indexing and more selective projections are future measured optimizations, not a claim
of constant-time activation.
What the benchmark measures
The engine, not a model. It drives the real HTTP and SSE paths with an instant scripted provider and an in-process echo environment, so no model latency reaches the numbers. First-token and turn figures include HTTP, request construction, committing to the session journal, and dispatch.
Density and reclaim measurements need Linux /proc/*/smaps_rollup. Other platforms run the
portable latency and correctness arms only. The harness refuses to substitute RSS for private
memory, because that would double-count shared pages. Any probe a subject cannot honestly answer
is recorded as a refusal rather than a number.
What CI enforces
The Linux HTTP integration checks two independently authenticated lazy providers, expiration without retry, restart history access without model execution, and repeated turns across suspended sessions. CI does not impose throughput or resident-memory thresholds during pre-launch iteration. Benchmark and leakage gates are deferred until representative workloads and baselines stabilize.
Correctness tests still bound journal growth on both axes. A transcript grows with every model call and
across every turn, so anything written per call or per turn would cost the sum of every intermediate
size rather than the final one; crates/brain/tests/journal_growth.rs holds the canonical journal
— and one public Event projection page — to a small constant multiple of the final context, and holds a
whole-transcript ratio so a session's log cannot grow with the square of its turn count.
History
Figures from before the current harness are archived in BENCHMARKS.md and should not be quoted as present-day performance.