Benchmark Results

July 24, 2026 · View on GitHub

User Memory Evaluation guide · Documentation index

This document provides a public result snapshot for OmniMemEval's current benchmark pipelines. The reproduced scores were generated under one evaluation harness so that memory backends are compared with the same data, prompts, answer model, judge model, and metric logic within each benchmark.

These results are intended to make comparison and reproduction easier. For each backend, the adapter and run configuration were prepared according to the product's public documentation, API reference, and available benchmark guidance. They are not a claim that every adapter has reached a globally optimal product-specific configuration. Contributions that improve an adapter's documented setup or default parameters are welcome.

Evaluation Setup

All reproduced runs used the same baseline evaluation configuration:

ComponentConfiguration
BenchmarksLoCoMo, LongMemEval, BEAM, PersonaMem v2, HaluMem
Memory service modelgpt-4.1-mini-2025-04-14 where the backend requires a model setting
Answer modelgpt-4.1-mini-2025-04-14
Judge modelgpt-4o-mini-2024-07-18
Primary metricLLM-as-a-judge accuracy for LoCoMo, LongMemEval, and HaluMem; nugget score for BEAM; rule-matching accuracy for PersonaMem v2
Efficiency metricAverage answer-stage context tokens

The reported Context Tokens value is the average number of tokens sent to the answer model per question, including the answer prompt and the retrieved context rendered by the memory backend. Lower context tokens indicate better token efficiency when accuracy is comparable.

Rows marked local/self-hosted were evaluated through a local or self-hosted service deployment because the managed cloud service was unavailable, insufficient for the full run, or not the recommended evaluation route at the time of testing.

Published reference scores are included only as external context. They may use different models, prompts, retrieval settings, context budgets, data versions, or judge implementations, and should not be treated as directly comparable to the reproduced OmniMemEval scores.

Result Summary

BackendLoCoMoLongMemEvalBEAM 100KBEAM 10MPersonaMem v2HaluMem
Mem077.6856.0070.4143.3336.7673.64
Zep / Graphiti63.8379.8068.7056.1132.3481.71
Supermemory73.5366.0765.9852.4939.6452.61
Viking69.3361.0770.7658.1430.8077.39
Cognee83.4851.8059.3056.0226.4672.60
Letta77.1277.6769.2252.3035.1285.43
Hindsight81.9972.2070.2259.7537.9883.99
Memori41.3420.80--33.1649.38
EverOS82.7580.4058.6447.7335.9488.66
MemMachine73.9063.6064.8051.9034.1447.02
mem973.6478.0065.7557.3030.7672.80
MemoryLake72.49-----
Backboard.io22.40-----
MemOS88.8389.2066.8756.7540.5880.91

A dash (-) means that a reproduced result is not included in this snapshot. For these missing cells, full runs were not completed under the same evaluation setup because of account/API access, service availability, benchmark support, or run-cost constraints. Partial or non-comparable runs are excluded rather than mixed into the reproduced result tables.

LoCoMo

LoCoMo evaluates long-conversation memory with multi-hop, temporal, and open-domain question answering. The reproduced evaluation excludes category 5 adversarial questions and covers 1,540 questions.

CategoryCountDescription
Single-Hop841Direct fact extraction from one evidence source
Multi-Hop282Reasoning over multiple conversation turns
Temporal321Time-aware retrieval and temporal reasoning
Open-Domain96Open-ended reasoning over multiple pieces of evidence

Reproduced Results

BackendDeploymentSingle-HopMulti-HopTemporalOpen-DomainOverallContext Tokens
Mem0cloud81.0976.1277.1554.1777.6817,395
Zepcloud65.3668.7955.5663.5463.831,862
Supermemorycloud75.3977.0767.6066.6773.5315,238
Vikingcloud78.0473.2948.8150.0069.335,964
Cogneecloud87.9978.8481.8363.1983.4832,532
Lettacloud87.9976.2453.4863.5477.1214,188
Hindsightcloud88.9878.8473.5258.3381.9924,683
Memoricloud47.3244.0922.5343.7541.348,139
EverOScloud86.8077.7884.1157.2982.758,559
MemMachinelocal/self-hosted83.4753.1971.9657.2973.902,577
mem9cloud79.2762.8873.6255.9073.641,597
MemoryLakecloud70.8775.3079.7554.1772.495,202
Backboard.iocloud25.0922.3413.4029.1722.401,198
MemOScloud92.5188.6585.0569.7988.835,400

Published Reference Results

BackendSingle-HopMulti-HopTemporalOpen-DomainOverallContext TokensSource
Mem094.695.492.582.392.56,956mem0.ai research
Zep96.494.095.679.294.75,760getzep research
Supermemory----77.1-Supermemory issue 795
Letta----74.0-Letta benchmark blog
Hindsight----92.0-Hindsight Benchmarks
Memori87.8772.7080.3763.5481.951,294Memori benchmark
EverOS96.6791.8489.7276.0493.05-EverMemOS paper
mem989.7183.1689.2564.5886.85-mem9
MemoryLake96.7991.8491.2885.4294.03-MemoryLake benchmark
Backboard.io89.3675.0091.9091.2090.00-Backboard LoCoMo repo

LongMemEval

LongMemEval evaluates long-term interactive memory across sessions. The OmniMemEval public pipeline uses the cleaned LongMemEval-S data by default.

CategoryCountDescription
single-session-user70User fact extraction from one historical session
single-session-assistant56Assistant-provided information extraction from one historical session
single-session-preference30User preference inference from one historical session
temporal-reasoning133Time-aware reasoning over session timestamps
multi-session133Reasoning over information from multiple sessions
knowledge-update78Selecting the latest valid answer after information changes

Reproduced Results

BackendDeploymentSS-UserSS-AsstSS-PrefTemp. ReasMulti-SKnow. UpdOverallContext Tokens
Mem0cloud8.5785.7196.6751.1350.3879.4956.00856
graphiti-zeplocal/self-hosted94.29100.0086.6774.4467.6779.4979.80117,106
Supermemorycloud87.1441.0765.5668.4260.1571.3766.076,635
Vikingcloud75.2446.4396.6755.3957.8960.2661.072,291
Cogneelocal/self-hosted67.1460.7183.3347.3737.5951.2851.8010,305
Lettacloud95.7198.2176.6769.4265.4182.0577.6749,431
Hindsightlocal/self-hosted82.8614.2996.6782.7171.4378.2172.2029,755
Memoricloud84.141.7923.333.7618.806.4120.802,779
EverOSlocal/self-hosted91.4389.2996.6781.9566.1779.4980.4012,379
MemMachinelocal/self-hosted75.7196.4383.3355.6439.8575.6463.602,803
mem9cloud95.7194.6456.6777.4462.4185.9078.003,805
MemOScloud100.00100.00100.0089.4778.9584.6289.204,151

Published Reference Results

BackendSS-UserSS-AsstSS-PrefTemp. ReasMulti-SKnow. UpdOverallContext TokensSource
Mem098.698.296.793.688.097.094.46,787mem0.ai research
Zep94.396.490.090.283.593.690.24,408getzep research
Supermemory97.0100.090.091.093.099.095.0-Supermemory LongMemBench
Hindsight------94.6-Hindsight Benchmarks
EverOS97.1485.7193.3377.4473.6889.7483.0-EverMemOS paper
Backboard.io97.198.290.091.791.793.693.4-Backboard LongMemEval repo

BEAM

BEAM evaluates long-term memory at different context scales. The public OmniMemEval runner supports 100K, 500K, 1M, and 10M scales; the reproduced snapshot below reports the 100K and 10M scales from the current result set.

BEAM uses nugget score rather than binary accuracy. Nugget score measures how well a generated answer covers the atomic reference facts, with 1.0 for fully covered, 0.5 for partially covered, and 0.0 for incorrect or missing evidence.

ScaleQuestionsDescription
100K400Baseline long-memory scale, approximately 128K tokens per conversation
10M200Extreme long-memory scale, approximately 10M tokens per conversation

Reproduced Results

BackendDeployment100K Nugget Score100K Context Tokens10M Nugget Score10M Context Tokens
Mem0cloud70.41 +/- 36.361,05543.33 +/- 41.991,086
graphiti-zeplocal/self-hosted68.70 +/- 36.468,66756.11 +/- 40.70176,211
Supermemorycloud65.98 +/- 37.106,29452.49 +/- 40.446,574
Vikingcloud70.76 +/- 35.232,02358.14 +/- 39.572,080
Cogneelocal/self-hosted59.30 +/- 39.4433,06556.02 +/- 41.3340,914
Lettacloud69.22 +/- 36.7682,01352.30 +/- 41.0649,786
Hindsightlocal/self-hosted70.22 +/- 36.4424,08559.75 +/- 39.6323,815
EverOSlocal/self-hosted58.64 +/- 39.367,39347.73 +/- 41.6611,657
MemMachinelocal/self-hosted64.80 +/- 37.625,41351.90 +/- 40.465,448
mem9cloud65.75 +/- 38.925,37257.30 +/- 39.034,947
MemOScloud66.87 +/- 37.371,63656.75 +/- 39.171,558

Published Reference Results

Backend100K Nugget Score100K Context Tokens10M Nugget Score10M Context TokensSource
Mem064.16,71948.66,914mem0.ai research
Hindsight75.0-64.1-Hindsight Benchmarks

PersonaMem v2

PersonaMem v2 evaluates personalized memory and preference-aware multiple-choice question answering. Accuracy is computed by matching the selected option against the gold answer; repeated runs are averaged.

Preference TypeCountDescription
ask_to_forget1,048Preferences that the user later asks the system to forget
neutral_preferences858Neutral personal preferences
anti_stereotypical_pref855Preferences that go against stereotypes
therapy_background627Therapy or mental-health background preferences
health_and_medical_conditions568Health and medical-condition preferences
stereotypical_pref533Stereotypical preference cases
sensitive_info511Sensitive personal information cases

Reproduced Results

BackendDeploymentAnti-StereotypicalAsk-to-ForgetHealth/MedicalNeutralSensitiveStereotypicalTherapyOverallContext Tokens
Mem0cloud41.4024.3338.3839.2830.7245.7843.5436.761,388
graphiti-zeplocal/self-hosted28.1929.7728.7031.9330.9237.9042.5832.343,645
Supermemorycloud42.3431.9742.9642.4230.7249.7240.6739.644,473
Vikingcloud31.5825.3829.2329.4933.6636.5934.7730.801,688
Cogneelocal/self-hosted18.8341.1316.9018.6532.8816.5134.9326.4610,189
Lettacloud31.4638.9332.3934.8527.4039.4039.2335.1230,903
Hindsightlocal/self-hosted33.5740.6536.6238.1136.0142.7838.1237.9815,926
Memoricloud30.1827.4832.3937.5333.4636.7738.1233.163,109
EverOScloud33.5731.4934.6837.7633.2744.4740.1935.946,572
MemMachinelocal/self-hosted28.1945.0432.2228.4429.7532.4638.6034.141,988
mem9cloud32.2822.9032.5734.5025.0538.2733.3330.762,045
MemOScloud33.8057.8232.5735.6636.5939.4039.2340.581,908

Published Reference Results

BackendOverallContext TokensSource
EverOS53.25-EverMemOS paper

HaluMem

HaluMem evaluates hallucination robustness in memory systems, including fact recall, boundary detection, conflicting memory handling, generalization, multi-hop inference, and dynamic updates.

Question TypeCountDescription
Basic Fact Recall746Extracting concrete facts from memory
Memory Boundary828Recognizing questions outside known memory rather than fabricating answers
Memory Conflict769Resolving conflicting memory records
Generalization & Application746Reasoning and applying known memories
Multi-hop Inference198Connecting multiple memories for inference
Dynamic Update180Answering with the latest valid state after memory updates

Reproduced Results

BackendDeploymentBasic FactBoundaryConflictGeneralizationMulti-HopDynamic UpdateOverallContext Tokens
Mem0cloud56.3091.9174.7778.4268.6942.2273.64803
graphiti-zeplocal/self-hosted69.4489.1388.4384.0579.8062.2281.715,404
Supermemorycloud25.2097.4644.8649.8741.9216.1152.611,672
Vikingcloud57.9192.5180.4984.7274.7547.7877.393,196
Cogneelocal/self-hosted52.4192.5174.6478.2869.7035.5672.608,981
Lettacloud85.9283.4585.7090.6284.3471.1185.4345,349
Hindsightlocal/self-hosted78.4287.4586.6186.1981.3173.8983.9914,798
Memoricloud20.9193.2441.3553.3524.2411.1149.383,275
EverOSlocal/self-hosted87.8088.8989.9990.4886.3680.5688.6610,824
MemMachinelocal/self-hosted20.3897.4234.7245.3126.7711.1147.021,093
mem9cloud54.5694.2068.7980.1670.7138.8972.80893
MemOScloud69.0391.3086.4883.5174.2455.0080.911,187

Published Reference Results

BackendOverallContext TokensSource
EverOS93.04-EverMind

Reproduction Notes

To reproduce a run, configure one of the templates under env_examples/, then run the corresponding benchmark script:

./scripts/run_locomo_eval.sh --lib memos --env .env.memos
./scripts/run_lme_eval.sh --lib memos --env .env.memos
./scripts/run_beam_eval.sh --lib memos --env .env.memos
./scripts/run_pmv2_eval.sh --lib memos --env .env.memos
./scripts/run_halumem_eval.sh --lib memos --env .env.memos

Replace memos with another adapter key to evaluate a different backend under the same benchmark pipeline. Use --version <name> to isolate result directories and make comparisons explicit.

Result artifacts are written under:

results/locomo/{LIB}-{VERSION}/
results/lme/{LIB}-{VERSION}/
results/beam/{LIB}-{VERSION}/
results/pmv2/{LIB}-{VERSION}/
results/halumem/{LIB}-{VERSION}/

Benchmark datasets are downloaded on demand and are not committed to this repository. See THIRD_PARTY_NOTICES.md for dataset license information.