PLATE · Ⅲ Benchmarks Open harness · reproducible

Honest benchmarks

The AI-memory market publishes leaderboard numbers that don't survive independent reproduction. We do it differently: every Mnem figure on this page comes from an open harness you can run yourself in one command — fully offline after the one-time dataset and embedding-model downloads. The only cloud component anywhere below is the benchmark's own judge model, and it is pinned.

Verified performance generated from harness output

LongMemEval end-to-end answer accuracy over all 500 questions — the figure the market quotes — with the official reader and judge prompts, a gpt-4o reader, and the benchmark's own judge pinned.
LongMemEval_s · QA ·
Retrieval Recall@5 on the same dataset: a labeled evidence session sits in Mnem's top-5 recalled sessions. No answering model, no judge — the layer we ship, measured alone.
LongMemEval_s · retrieval · questions
Retrieval Recall@10 — the context window the end-to-end reader actually receives. The gap to 100% is the ceiling on the QA number above.
LongMemEval_s · retrieval · top-10
On-device recall latency at p95, measured through the shipped search path with no network round-trip. The column that decides product fit.
Local harness · recall p95 ·
† Method notes — what was measured, and how

Our rules

  1. Loading…

Mnem, measured locally open harness

CategoryItemsRecall@1Recall@5
Loading results…

Reproduce it: python benchmarks/run_benchmarks.py — dataset, scorer, and config ship with the app's benchmark kit. Offline after the one-time embedding-model download (the harness aborts rather than silently substituting the non-semantic fallback encoder).

What competitors claim vs. what reproductions found

These are not our measurements. Claims are the vendors' own self-reported figures; reproductions are third-party runs (some by competing vendors — noted where known). We list both because, as of 2026, cross-vendor leaderboard numbers are harness-relative and not comparable. The one axis they share with us is the end-to-end LongMemEval figure above — read the caveats there first.

The columns that decide product fit

Where memories liveRecall pathPortability

Competitor claims and reproductions cited as of the date in each entry. Nothing on this page is hand-edited: figures are generated from harness output.

Run the numbers yourself.

Sub-10ms recall on your machine, with nothing leaving it.