Aex Brain
Reference

Benchmarks

Performance diagnostics and CI correctness coverage.

The current regression suite proves lazy history access, turn-end passivation, isolated subscription backlogs, and native Tool dispatch at full parent concurrency. Previous headline numbers predate this design and are archived in BENCHMARKS.md. They are not release thresholds for the current implementation.

Measure cold readiness, explicit admission, prepared creation, first activation, resumed activation, and history reads separately. Include both the server and Wasmtime worker in memory and CPU totals. Use retained sessions with long histories; an empty-server RSS smoke cannot establish session density.

The existing journal probe reports canonical reopen, disposable checkpoint write, and cached reopen on the same workload. Run cargo test --release -p brain --test journal_throughput -- --ignored --nocapture. A checkpoint saves decoding canonical history, but its index and current transcript still grow with history. Disk indexing and more selective projections are future measured optimizations, not a claim of constant-time activation.

What the benchmark measures

The engine, not a model. It drives the real HTTP and SSE paths with an instant scripted provider and an in-process echo environment, so no model latency reaches the numbers. First-token and turn figures include HTTP, request construction, committing to the session journal, and dispatch.

Density and reclaim measurements need Linux /proc/*/smaps_rollup. Other platforms run the portable latency and correctness arms only. The harness refuses to substitute RSS for private memory, because that would double-count shared pages. Any probe a subject cannot honestly answer is recorded as a refusal rather than a number.

What CI enforces

The Linux HTTP integration checks two independently authenticated lazy providers, expiration without retry, restart history access without model execution, and repeated turns across suspended sessions. CI does not impose throughput or resident-memory thresholds during pre-launch iteration. Benchmark and leakage gates are deferred until representative workloads and baselines stabilize.

Correctness tests still bound journal growth on both axes. A transcript grows with every model call and across every turn, so anything written per call or per turn would cost the sum of every intermediate size rather than the final one; crates/brain/tests/journal_growth.rs holds the canonical journal — and one public Event projection page — to a small constant multiple of the final context, and holds a whole-transcript ratio so a session's log cannot grow with the square of its turn count.

History

Figures from before the current harness are archived in BENCHMARKS.md and should not be quoted as present-day performance.

On this page