EVALUATION DIRECTORY
Benchmarks & Datasets
A consolidated repository of evaluations and environments referenced in the survey (Table 8). Filter by memory capabilities or environmental structures.
| Name | Fac. | Exp. | M.M. | Environment | Scale / Size | Primary Features | Link |
|---|---|---|---|---|---|---|---|
| MemBench | simulated | 53,000 samples | interactive scenarios | Link | |||
| MemoryAgentBench | simulated | 4 tasks | multi-turn interactions | Link | |||
| LoCoMo | real | 300 samples | conversational memory | Link | |||
| WebChoreArena | real | 4 tasks / 532 samples | tedious web browsing | Link | |||
| MT-Mind2Web | real | 720 samples | conversational web navigation | Link | |||
| PersonaMem | simulated | 15 tasks / 180 samples | dynamic user profiling | Link | |||
| LongMemEval | simulated | 5 tasks / 500 samples | interactive memory | Link | |||
| PerLTQA | simulated | 8,593 samples | social personalized interactions | Link | |||
| MemoryBank | simulated | 194 samples | user memory updating | Link | |||
| MPR | simulated | 108,000 samples | user personalization | Link | |||
| PrefEval | simulated | 3,000 samples | personal preferences | Link | |||
| LOCCO | simulated | 3,080 samples | chronological conversations | Link | |||
| StoryBench | mixed | 3 tasks | interactive fiction games | Link | |||
| MemoryBench | simulated | 4 tasks / ~20,000 samples | continual learning | Link | |||
| Madial-Bench | simulated | 331 samples | memory recalling | Link | |||
| Evo-Memory | simulated | 10 tasks / ~3,700 samples | test-time learning | Link | |||
| LifelongAgentBench | simulated | 1,396 samples | lifelong learning | Link | |||
| StreamBench | simulated | 9,702 samples | continuous online learning | Link | |||
| DialSim | real | ~1,300 samples | multi-dialogue understanding | Link | |||
| LongBench | mixed | 21 tasks / 4,750 samples | long-context understanding | Link | |||
| LongBench v2 | mixed | 20 tasks / 503 samples | long-context multitasks | Link | |||
| RULER | simulated | 13 tasks | long-context retrieval | Link | |||
| BABILong | simulated | 20 tasks | long-context reasoning | Link | |||
| MM-Needle | simulated | ~280,000 samples | multimodal long-context retrieval | Link | |||
| HaluMem | simulated | 3,467 samples | memory hallucinations | Link | |||
| HotpotQA | simulated | 113k samples | long-context QA | Link | |||
| ALFWorld | simulated | 3,353 tasks | text-based embodied environment | Link | |||
| ScienceWorld | simulated | 10 t. / 30 t. | interactive embodied environment | Link | |||
| AgentGym | mixed | 89 tasks / 20,509 samples | multiple environments | Link | |||
| AgentBoard | mixed | 9 tasks / 1013 samples | multi-round interaction | Link | |||
| WebShop | simulated | 12,087 samples | e-commerce web interaction | Link | |||
| WebArena | real | 812 samples | web interaction | Link | |||
| MMInA | real | 1,050 samples | multihop web interaction | Link | |||
| SWE-Bench Verified | real | 500 samples | code repair | Link | |||
| GAIA | real | 466 samples | human-level deep research | Link | |||
| xBench-DS | real | 100 samples | deep-search evaluation | Link | |||
| ToolBench | real | 126,486 samples | API tool use | Link | |||
| GenAI-Bench | real | ~40,000 samples | visual generation evaluation | Link |