Bench results
One result on this page carries full provenance, and it is the first section below. Everything else is either published with its defect stated, or refused with the reason given. That is a shorter page than most vendors would write and it is the honest length.
This page is reference. It is the single place benchmark numbers for the system are published.
1. The house rule
A number carries the corpus it was measured on and the revision it was measured at, or it is not quoted. Not softened, not hedged, dropped.
The reason is specific rather than pious. LongMemEval-S and LongMemEval-M give materially different numbers on the same code, and quoting a lift measured on one as though it were measured on the other is the failure mode this rule exists to prevent. A recall figure without its corpus is not a claim anyone can check, including us.
One caveat applies to everything below. CoreCrux is a closed system, so a revision identifier here is an attestation, not something you can independently check. That is unlike the Crux Daemon claims elsewhere in this documentation, where every citation links to a line of source you can read. Treat these as you would treat any vendor's figures: useful, dated, specific enough to argue with, and weaker evidence than anything you can verify yourself. If you are evaluating us, ask for the run record.
None of these is a service level. They are measurements of dated runs on stated hardware. No latency, throughput or availability commitment attaches to any of them, and you should not size a deployment from a figure you did not measure on your own.
2. Retrieval quality: LongMemEval, 40-problem cohort
| Corpus | LME-40 |
| Date | 2026-05-29 |
| Revision | 89d965f (CoreCrux PR #82) |
LME-40 is a 40-problem mixed cohort drawn from LongMemEval: 20 LME-S robust-failure problems and the first 20 LME-M problems, run as paired tenants, one freshly ingested and one on the earlier baseline ingest. 80 probes. Same endpoint, same question and same anchor on both sides of each pair.
| Metric | Baseline ingest | Fresh ingest | Delta |
|---|---|---|---|
| Recall@1 | 60 % | 78 % | +18 pp |
| MRR | 0.717 | 0.806 | +0.089 |
| Recall@10 | 88 % | 88 % | no change |
Read the third row as carefully as the first two. Recall@10 ties. The whole of the improvement is in where the right evidence lands in the ranking, not in whether it is retrieved at all. If your application shows a user ten results, this result predicts no change for you.
2.1 The decomposition
Publishing the split matters, because otherwise the whole gain gets attributed to one thing.
| Contribution | Recall@1 | Basis |
|---|---|---|
| Clean ingest base alone | 47 % to 80 %, +33 pp | The 15 tenant pairs on which the additional lane did not participate |
| The additional retrieval lane | 68 % to 76 %, +8 pp | The 25 tenant pairs on which it did |
Most of the gain is corpus hygiene. A smaller, real gain is the additional lane. Anyone quoting +18 pp as the effect of a retrieval feature is quoting it wrong, and that includes us.
The full write-up, with the negative result that sits beside it, is in CoreCrux §2.7.
3. Results we hold and do not publish here
Three, and the reasons differ.
Two results with no revision recorded. A substrate-coverage measurement on LME-M and a production integration throughput comparison both exist, both name their corpus and workload, and neither has a revision in its run record. That is a defect in how those runs were recorded. They are published in CoreCrux §2.7 with that defect stated on the page, and they are not repeated here as headline numbers, because this page is the one that holds the line. If you want only fully-provenanced figures, use section 2 and ignore them.
A figure on the full 500-problem LongMemEval set. There is a widely quoted internal number for this. It is not published and it will not be, because it was banked through a gated-tail splice rather than a clean run, and there is no full-500 result on the live backend to replace it with. A spliced result is not a result. If you have seen that figure quoted at you by anyone, including by us in an earlier document, treat it as withdrawn.
A strict-mode recollection from the VaultCrux build. The founder's account of that period, on the opening page, mentions returning over 90 % on the bench in strict mode. It is a recollection with no corpus name and no revision attached, so it does not become a benchmark claim by being written down. It is kept in the narrative because it is what happened, and it is disclaimed there for the same reason it is disclaimed here.
This section is the most useful part of the page. A vendor's benchmark table tells you what they measured; what they refuse to publish tells you what their numbers are worth.
4. Corpus glossary
Use these names whenever you quote one of our numbers back at us.
| Name | What it is |
|---|---|
| LME-S | LongMemEval small. Roughly 250 events per problem, chat-history shaped. |
| LME-M | LongMemEval medium. Roughly 2,500 events per problem. |
| LME-40 | The 40-problem mixed cohort: 20 LME-S robust-failure problems plus the first 20 LME-M, run as paired fresh-ingest and baseline-ingest tenants. |
LME-S, LME-M and LME-40 give materially different numbers. A recall figure quoted without saying which of the three it came from is not a claim we will stand behind, including when we are the ones who said it.
5. What a benchmark result is not
A retrieval score is evidence about retrieval quality on one corpus at one revision. It is not evidence that an answer built on that retrieval is correct, and it is not evidence about what an agent did with it. Those are separate questions with separate machinery: see Receipts and proof for what a receipt establishes, and Assurance and compliance for the controls that are built but not enforced.
6. Read next
- CoreCrux §2.7, the full measured-results section including the negative result.
- Retrieval and budgets, what the free local daemon's retrieval actually does, with source links.
- Assurance and compliance, the dated disclosure.

