Measurements
Dated product proof and its dataset tags.
Source: docs/initiatives/memory-api/00_MEASUREMENTS.md — this page is rendered from it, not copied. Where the home page's board numbers come from, with the dataset tag and the command that reproduces each one.
Date: 2026-07-26 · Status: measured, reproducible · Dataset tag: real founder corpus, Cloud SQL brain_rebuild, user founder, live spine (events 42,081 · memories 2,365 · search_chunks 11,216 · embeddings 10,643 · kg_entities 408 · people 1,460 · fact_assertions 7,042).
Gates every number in the API design. Two of the three planned experiments ran; the third (contextual retrieval) is scoped at the end and has not run.
Reproduce:
.venv/bin/python scripts/measure_fast_context.py --out artifacts/FAST_CONTEXT_2026_07_26.jsonArtifact: artifacts/FAST_CONTEXT_2026_07_26.json (per-case grades, windows, latencies).
1. The headline: generation is currently SUBTRACTIVE
The question was whether a fast Context — retrieved, cited evidence that composes nothing — still carries the answer.
CORRECTED 2026-07-29. This section said "zero generative calls" and the table below recorded
0. Both were false when written:_envelopereturned a hardcoded0whilebrain/retrieve/rewrite.py:expand()called Gemini on every request. The real cost was 2 — one embedding, one generative rewrite. The rewrite has since been removed from thefastpath (artifacts/_ablation_fast_rewrite/, branch DROP), sofastnow spends 1 model call: the query embedding, which the ledger charges. The carry figures below are unaffected — they were graded on retrieved windows, and the rewrite did not move them (clean arm 1-1, p=0.75). Only the COST column was wrong.AND THEN THE RECEIPT SPLIT, 2026-07-29 (later the same day). For a few hours this note said that one call was what
cost.generative_callsreports. It is not, and that was the same category error one layer up: an embedding is not a generative call.CostReportnow carries three fields, and a readyfastContext reportsmodel_calls_total: 1 · embedding_calls: 1 · generative_calls: 0. Read the column below as model calls, which is what it counts and what the header now says.Graded on the same 43 held-out founder-confirmed cases, with the ruler's
own deterministic crit-fact grader (brain.eval.heldout.grader; writer ≠ checker holds, B4).
| Path | Model calls AS RUN (see note) | Carries / serves the answer | Confidently wrong |
|---|---|---|---|
| Non-generative Context, window 12 | 2 | 32/43 (74%) | 0 — structural |
| Non-generative Context, window 8 | 2 | 30/43 (70%) | 0 — structural |
Full serve path ask() (v2-40 ruler, 2026-07-24) | up to 6 | 27/43 (63%) | 0 |
The evidence already contains five answers the generative path fails to serve. Route → synthesize → judge is not adding comprehension on this corpus; it is losing cases retrieval found.
And the fast tier cannot be confidently wrong, because it asserts nothing. It hands over cited evidence for the caller's model to read. It can be incomplete; it cannot fabricate. That is a structural property, not a score — which is exactly the kind of guarantee a memory API can sell.
Trap behaviour: of 8 traps, the bait text appeared in the evidence window in 1 (V2-04). That is a retrieval observation, not a B14 safety failure — conflating them would launder a retrieval fact into a safety fact, which the grader doctrine forbids.
Context size is comfortable: median 2,585 estimated tokens at window 12 (max 8,633), well inside a 4k budget.
The evidence-window budget costs 2 cases
Window 8 → 30/43; window 12 → 32/43. Two cases have their answer retrieved but crowded out of the synthesizer's evidence window. This quantifies the "serve-window crowding class" already named in status/BRAIN_STATE_OF_THE_UNION_2026_07_24.md — it is worth 2 cases, and budget is the knob that exposes the trade to the caller.
2. The latency finding: generation was never the bottleneck
The design assumed removing generation was the route to the 200 ms SLO. It is not. Retrieval alone measured p50 2,437 ms · p95 3,495 ms over the 43 cases — and note those numbers INCLUDE the generative query rewrite nobody had noticed was on the path. With it removed (2026-07-29): p50 1,433.5 ms · p95 2,163.5 ms · p99 11,073.0 ms, n=258 → see the run note. Still not 200 ms, which was the finding then and remains it now.
NAME THE RUN, 2026-07-29. Two ablations ran that day and this line quoted the first one without saying so, which reads as a contradiction of
docs/status/CURRENT_TRUTH.md§1b:
artifact n (variant B, counted, error-free) p50 p95 p99 _ablation_fast_rewrite/20260729T024204Z_all.json(quoted above)257 1,433.5 2,163.5 11,073.0 _ablation_fast_rewrite/20260729T031018Z_all.json(the quotable one,CURRENT_TRUTH.md§1b)258 1,433.4 2,085.0 11,116.2 Both recomputed here from the raw artifacts with the same nearest-rank rule the harness uses (
index = ceil(q·n) − 1, no interpolation,scripts/_ablation_fast_rewrite.py:448). The two agree on p50 to 0.1 ms and disagree on p95 by 79 ms — sampling, not a change in the system, and the conclusion (not 200 ms; p95 above 2 s either way) is identical. Quote §1b ofCURRENT_TRUTH.mdwhen a number leaves this repo, and always carry the run id: an unlabelled p95 is the reason this note exists.
Breakdown (12 queries, pooled connection and transport, embedding precomputed where noted):
| Component | p50 | Note |
|---|---|---|
embed_query (Vertex round trip) | 170 ms | p95 181 ms. Not the problem, and cacheable — agents repeat queries. |
| Lanes + fusion | 1,586 ms | The bottleneck. 90% of a pooled request. |
| Per-call setup (fresh DB conn + GeminiClient) | 464 ms | Pure waste; a pooled server pays zero. The measurement harness paid it. |
Per-lane, and why this is good news
| Lane | p50 | max |
|---|---|---|
state_lane | 425 ms | 804 ms |
graph_lane | 352 ms | 539 ms |
fts_lane | 347 ms | 1,043 ms |
person_lane | 97 ms | 286 ms |
vector_lane | 90 ms | 191 ms |
temporal_lane | 65 ms | 418 ms |
life_event_lane | 40 ms | 159 ms |
episode_lane | 34 ms | 37 ms |
calendar_lane | 33 ms | 65 ms |
kinship_lane | 0 ms | 0 ms |
| sum of medians | 1,483 ms | ≈ the 1,586 ms measured — the lanes run sequentially |
The slowest single lane is 425 ms against a sequential sum of 1,483 ms. So the lanes were made to run concurrently, each on its own pooled read connection (brain/spine/pool.py), with the embed submitted into the same fan-out so its 170 ms overlaps the eight lanes that do not need meaning.
Fan-out: built and measured, 2026-07-26
| p50 | p95 | max | |
|---|---|---|---|
| Sequential (pooled conn+transport) | 1,756 ms | 2,129 ms | 2,255 ms |
| Concurrent, warm | 1,020 ms | 1,078 ms | 1,176 ms |
| End-to-end over the 43 cases (harness pays fresh setup) | 2,437 → 1,465 ms | 3,495 → 1,872 ms |
Measured speedup 1.72× — not the ~3.5× predicted above. The prediction assumed the lanes were independent; they are independent in code but they contend for one Postgres instance. Wall clock did not fall to the slowest lane (425 ms) because under concurrency every lane gets slower. The database is the shared resource, and that is the honest ceiling on this technique.
Correctness: proven identical. Fused candidate order compared sequential vs concurrent across all 43 cases — 43/43 byte-identical. The non-generative board is unchanged (30/43 at window 8, 32/43 at window 12, same single trap-bait case). It is a scheduling change and nothing else.
Consequence for the roadmap: the remaining path to 200 ms is not more concurrency. It is making the three slow lanes cheaper — state (425 ms), graph (352 ms), fts (347 ms) — a targeted indexing problem, plus connection reuse (pg_connect measures 320 ms, so a server must never open one per request).
A correction to the design brief
The vector lane costs 90 ms. ANN over 10,643 halfvec(3072) vectors is not a latency problem, so the embedding-dimension decision is purely a storage and index-memory question at scale — it must not be justified on serving latency. The design brief implied an ANN-latency motive; that is wrong and is corrected there.
3. What this means for the product
fastis the product. It carries more (74% vs 63%) for one model call against up to six, and cannot fabricate.thoroughis the upsell, and it must be shown to add something before it is worth its price — on this corpus it currently subtracts. > CORRECTED 2026-07-29 — the cost RATIO is withdrawn; the call counts replace it. This read > "costs ~1000× less (one embedding call versus six generative calls at ~$0.0075)". Two things > were wrong with it. The numerator was not one call — when this was writtenfastwas > spending two, an embedding plus an undetected generative query rewrite, so the comparison was > 2-vs-6 dressed as 1-vs-6. And the denominator of the ratio was a zero that is not zero: the > embedding is a real Vertex round trip (p50 170 ms, §2), never priced here, so "1000× less" was > a generation-token ratio presented as a request-cost ratio. The$0.0075is a real figure but > it is theask()path's per-answer cost fromdocs/status/STRESS_REPORT_2026_07.md, not a > per-Context one. 1 model call vs up to 6 is countable, verifiable and enough — a priced > ratio needs the embedding's price, and nobody has looked it up.- The 200 ms SLO is credible but not free, and the work is lane concurrency plus three SQL indexes — not model choice, not a rewrite, not a second store.
- Publish
p50/p95next toconfidently_wrong_rate. Both are now measured.
4. Two live defects found while connecting (not part of the experiment)
Both concern BRAIN_DATABASE_URL and both are real today:
db.ping()anddb.table_exists()return False against a perfectly healthy database. Thecom.brain.synclaunchd agent supplies a barepostgresql://…URL. SQLAlchemy resolves that to psycopg2, which is not installed (ModuleNotFoundError: No module named 'psycopg2'), so every SQLAlchemy-path helper fails closed./healthreports the DB down while the pipeline runs fine, and every liveness detector built ontable_exists()would conclude no table exists.- No single value of the variable satisfies both paths.
brain/spine/db.py:70strips+psycopgfor the raw psycopg3 path, andfounder_confirmed_v2._data_now()callspsycopg.connect(os.environ["BRAIN_DATABASE_URL"])directly — which requires the bare form — whiledb.ping()requires the+psycopgform. The two are mutually exclusive.
Fix: normalize in one place (config.load_settings), so the SQLAlchemy URL always carries +psycopg and the raw path always strips it, and route _data_now() through db.pg_connect(). One-line-class changes, but they need the gate, so they are filed here rather than made.
5. Not run
Experiment 3 — contextual retrieval. Prepending a short statement of where a chunk sits before embedding it, published at ~35% fewer top-20 retrieval failures (~49% with a lexical lane, which we have). Requires re-embedding a bounded sample and re-running this same harness — the harness now exists, so this is a scoped, cheap follow-up. Given §1's result (evidence carries 74%, the misses are retrieval misses, not synthesis misses), this is now the highest-value remaining experiment: every one of the 11 misses at window 12 is a case where the answer was not retrieved at all.
# 00 — Measurements: the non-generative Context, and where the latency actually is **Date:** 2026-07-26 · **Status:** measured, reproducible · **Dataset tag:** real founder corpus, Cloud SQL `brain_rebuild`, user `founder`, live spine (events 42,081 · memories 2,365 · search_chunks 11,216 · embeddings 10,643 · kg_entities 408 · people 1,460 · fact_assertions 7,042). Gates every number in the API design. Two of the three planned experiments ran; the third (contextual retrieval) is scoped at the end and has not run. Reproduce: ```bash .venv/bin/python scripts/measure_fast_context.py --out artifacts/FAST_CONTEXT_2026_07_26.json ``` Artifact: `artifacts/FAST_CONTEXT_2026_07_26.json` (per-case grades, windows, latencies). --- ## 1. The headline: generation is currently SUBTRACTIVE The question was whether a `fast` Context — retrieved, cited evidence that composes nothing — still carries the answer. > **CORRECTED 2026-07-29.** This section said "zero generative calls" and the table below recorded > `0`. Both were false when written: `_envelope` returned a hardcoded `0` while > `brain/retrieve/rewrite.py:expand()` called Gemini on every request. The real cost was **2** — one > embedding, one generative rewrite. The rewrite has since been removed from the `fast` path > (`artifacts/_ablation_fast_rewrite/`, branch DROP), so `fast` now spends **1** model call: the > query embedding, which the ledger charges. The carry figures below > are unaffected — they were graded on retrieved windows, and the rewrite did not move them (clean > arm 1-1, p=0.75). Only the COST column was wrong. > > **AND THEN THE RECEIPT SPLIT, 2026-07-29 (later the same day).** For a few hours this note said > that one call was what `cost.generative_calls` reports. It is not, and that was the same category > error one layer up: **an embedding is not a generative call.** `CostReport` now carries three > fields, and a ready `fast` Context reports `model_calls_total: 1 · embedding_calls: 1 · > generative_calls: 0`. Read the column below as *model calls*, which is what it counts and what > the header now says. > > Graded on the same 43 held-out founder-confirmed cases, with the ruler's own deterministic crit-fact grader (`brain.eval.heldout.grader`; writer ≠ checker holds, B4). | Path | Model calls AS RUN (see note) | Carries / serves the answer | Confidently wrong | |---|---:|---:|---:| | **Non-generative Context, window 12** | **2** | **32/43 (74%)** | **0 — structural** | | Non-generative Context, window 8 | 2 | 30/43 (70%) | 0 — structural | | Full serve path `ask()` (v2-40 ruler, 2026-07-24) | up to 6 | 27/43 (63%) | 0 | **The evidence already contains five answers the generative path fails to serve.** Route → synthesize → judge is not adding comprehension on this corpus; it is losing cases retrieval found. **And the fast tier cannot be confidently wrong, because it asserts nothing.** It hands over cited evidence for the caller's model to read. It can be incomplete; it cannot fabricate. That is a structural property, not a score — which is exactly the kind of guarantee a memory API can sell. Trap behaviour: of 8 traps, the bait text appeared in the evidence window in **1** (V2-04). That is a retrieval observation, not a B14 safety failure — conflating them would launder a retrieval fact into a safety fact, which the grader doctrine forbids. Context size is comfortable: median **2,585 estimated tokens** at window 12 (max 8,633), well inside a 4k budget. ### The evidence-window budget costs 2 cases Window 8 → 30/43; window 12 → 32/43. Two cases have their answer retrieved but crowded out of the synthesizer's evidence window. This quantifies the "serve-window crowding class" already named in `status/BRAIN_STATE_OF_THE_UNION_2026_07_24.md` — it is worth 2 cases, and `budget` is the knob that exposes the trade to the caller. --- ## 2. The latency finding: generation was never the bottleneck The design assumed removing generation was the route to the 200 ms SLO. **It is not.** Retrieval alone measured **p50 2,437 ms · p95 3,495 ms** over the 43 cases — and note those numbers INCLUDE the generative query rewrite nobody had noticed was on the path. With it removed (2026-07-29): **p50 1,433.5 ms · p95 2,163.5 ms · p99 11,073.0 ms**, n=258 → see the run note. Still not 200 ms, which was the finding then and remains it now. > **NAME THE RUN, 2026-07-29.** Two ablations ran that day and this line quoted the first one > without saying so, which reads as a contradiction of `docs/status/CURRENT_TRUTH.md` §1b: > > | artifact | n (variant B, counted, error-free) | p50 | p95 | p99 | > |---|---:|---:|---:|---:| > | `_ablation_fast_rewrite/20260729T024204Z_all.json` (quoted above) | 257 | 1,433.5 | 2,163.5 | 11,073.0 | > | `_ablation_fast_rewrite/20260729T031018Z_all.json` (**the quotable one**, `CURRENT_TRUTH.md` §1b) | 258 | 1,433.4 | 2,085.0 | 11,116.2 | > > Both recomputed here from the raw artifacts with the same nearest-rank rule the harness uses > (`index = ceil(q·n) − 1`, no interpolation, `scripts/_ablation_fast_rewrite.py:448`). The two > agree on p50 to 0.1 ms and disagree on p95 by 79 ms — sampling, not a change in the system, and > the conclusion (*not 200 ms; p95 above 2 s either way*) is identical. **Quote §1b of > `CURRENT_TRUTH.md` when a number leaves this repo**, and always carry the run id: an unlabelled > p95 is the reason this note exists. Breakdown (12 queries, pooled connection and transport, embedding precomputed where noted): | Component | p50 | Note | |---|---:|---| | `embed_query` (Vertex round trip) | **170 ms** | p95 181 ms. Not the problem, and cacheable — agents repeat queries. | | **Lanes + fusion** | **1,586 ms** | **The bottleneck.** 90% of a pooled request. | | Per-call setup (fresh DB conn + GeminiClient) | 464 ms | Pure waste; a pooled server pays zero. The measurement harness paid it. | ### Per-lane, and why this is good news | Lane | p50 | max | |---|---:|---:| | `state_lane` | **425 ms** | 804 ms | | `graph_lane` | **352 ms** | 539 ms | | `fts_lane` | **347 ms** | 1,043 ms | | `person_lane` | 97 ms | 286 ms | | `vector_lane` | **90 ms** | 191 ms | | `temporal_lane` | 65 ms | 418 ms | | `life_event_lane` | 40 ms | 159 ms | | `episode_lane` | 34 ms | 37 ms | | `calendar_lane` | 33 ms | 65 ms | | `kinship_lane` | 0 ms | 0 ms | | **sum of medians** | **1,483 ms** | ≈ the 1,586 ms measured — **the lanes run sequentially** | **The slowest single lane is 425 ms against a sequential sum of 1,483 ms.** So the lanes were made to run concurrently, each on its own pooled read connection (`brain/spine/pool.py`), with the embed submitted into the same fan-out so its 170 ms overlaps the eight lanes that do not need meaning. ### Fan-out: built and measured, 2026-07-26 | | p50 | p95 | max | |---|---:|---:|---:| | Sequential (pooled conn+transport) | 1,756 ms | 2,129 ms | 2,255 ms | | **Concurrent, warm** | **1,020 ms** | **1,078 ms** | 1,176 ms | | End-to-end over the 43 cases (harness pays fresh setup) | 2,437 → **1,465 ms** | 3,495 → **1,872 ms** | | **Measured speedup 1.72× — not the ~3.5× predicted above.** The prediction assumed the lanes were independent; they are independent in *code* but they contend for one Postgres instance. Wall clock did not fall to the slowest lane (425 ms) because under concurrency every lane gets slower. **The database is the shared resource, and that is the honest ceiling on this technique.** **Correctness: proven identical.** Fused candidate order compared sequential vs concurrent across all 43 cases — **43/43 byte-identical**. The non-generative board is unchanged (30/43 at window 8, 32/43 at window 12, same single trap-bait case). It is a scheduling change and nothing else. **Consequence for the roadmap:** the remaining path to 200 ms is **not** more concurrency. It is making the three slow lanes cheaper — `state` (425 ms), `graph` (352 ms), `fts` (347 ms) — a targeted indexing problem, plus connection reuse (`pg_connect` measures **320 ms**, so a server must never open one per request). ### A correction to the design brief **The vector lane costs 90 ms.** ANN over 10,643 halfvec(3072) vectors is not a latency problem, so the embedding-dimension decision is **purely a storage and index-memory question at scale** — it must not be justified on serving latency. The design brief implied an ANN-latency motive; that is wrong and is corrected there. --- ## 3. What this means for the product 1. **`fast` is the product.** It carries more (74% vs 63%) for **one model call against up to six**, and cannot fabricate. `thorough` is the upsell, and it must be shown to *add* something before it is worth its price — on this corpus it currently subtracts. > **CORRECTED 2026-07-29 — the cost RATIO is withdrawn; the call counts replace it.** This read > *"costs ~1000× less (one embedding call versus six generative calls at ~$0.0075)"*. Two things > were wrong with it. **The numerator was not one call** — when this was written `fast` was > spending two, an embedding plus an undetected generative query rewrite, so the comparison was > 2-vs-6 dressed as 1-vs-6. **And the denominator of the ratio was a zero that is not zero**: the > embedding is a real Vertex round trip (p50 170 ms, §2), never priced here, so "1000× less" was > a generation-token ratio presented as a request-cost ratio. The `$0.0075` is a real figure but > it is the *`ask()` path's* per-answer cost from `docs/status/STRESS_REPORT_2026_07.md`, not a > per-Context one. **1 model call vs up to 6 is countable, verifiable and enough** — a priced > ratio needs the embedding's price, and nobody has looked it up. 2. **The 200 ms SLO is credible but not free**, and the work is lane concurrency plus three SQL indexes — not model choice, not a rewrite, not a second store. 3. **Publish `p50/p95` next to `confidently_wrong_rate`.** Both are now measured. --- ## 4. Two live defects found while connecting (not part of the experiment) Both concern `BRAIN_DATABASE_URL` and both are real today: 1. **`db.ping()` and `db.table_exists()` return False against a perfectly healthy database.** The `com.brain.sync` launchd agent supplies a bare `postgresql://…` URL. SQLAlchemy resolves that to psycopg2, which is not installed (`ModuleNotFoundError: No module named 'psycopg2'`), so every SQLAlchemy-path helper fails closed. `/health` reports the DB down while the pipeline runs fine, and every liveness detector built on `table_exists()` would conclude no table exists. 2. **No single value of the variable satisfies both paths.** `brain/spine/db.py:70` *strips* `+psycopg` for the raw psycopg3 path, and `founder_confirmed_v2._data_now()` calls `psycopg.connect(os.environ["BRAIN_DATABASE_URL"])` directly — which requires the bare form — while `db.ping()` requires the `+psycopg` form. The two are mutually exclusive. **Fix:** normalize in one place (`config.load_settings`), so the SQLAlchemy URL always carries `+psycopg` and the raw path always strips it, and route `_data_now()` through `db.pg_connect()`. One-line-class changes, but they need the gate, so they are filed here rather than made. --- ## 5. Not run **Experiment 3 — contextual retrieval.** Prepending a short statement of where a chunk sits before embedding it, published at ~35% fewer top-20 retrieval failures (~49% with a lexical lane, which we have). Requires re-embedding a bounded sample and re-running this same harness — the harness now exists, so this is a scoped, cheap follow-up. Given §1's result (evidence carries 74%, the misses are retrieval misses, not synthesis misses), **this is now the highest-value remaining experiment**: every one of the 11 misses at window 12 is a case where the answer was not retrieved at all.