---
format: "aidr-story-markdown/v1"
id: "b4632004133b61fc35d40b5d8af7d4fd62ae18718dda8687bc8c1d881810f28b"
canonical_url: "https://aidr.today/b4632004?lang=en"
title: "Agent Zero Memory Cuts LLM Query Costs 30x, New Benchmark High of 95.60%"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-02T08:57:34.000Z"
category: "Research"
topics: ["agent","llm","benchmark","inference"]
source_urls: ["https://huggingnews.com/ai/agent-zero-memory-cuts-llm-query-costs-30x-new-benchmark-high-of-9560per-50c0feb9","https://x.com/dair_ai/status/2094953486047977860","https://x.com/mrdj1968/status/2094956976161878121","https://x.com/HEI/status/2094767863844237527"]
summary: "Ming Wu and Pengyuan Zhu's Agent Zero Memory architecture improves long-term recall for LLM agents by separating memory into three concurrent search systems. The framework runs an episodic events timeline, an entity-event knowledge graph, and a curated documentary memory of durable facts to reach 95.60% on LongMemEval and 93.60% on LoCoMo. Evaluation of eight backbone models found a 3.4 point accuracy increase and a 30x reduction in per-query costs. Retrieval starts with an intent gate and a source router to narrow search buckets. The \"citation lock\" mechanism ensures replies only cite evidence that a reader has actually opened, while indexed raw sources prevent data loss during bad extractions. The authors have not yet performed ablation studies to determine if the accuracy gains result from the memory stores or the intent gate."
---

# Agent Zero Memory Cuts LLM Query Costs 30x, New Benchmark High of 95\.60%

> [Open the canonical story](<https://aidr.today/b4632004?lang=en>)

**Published:** 2026-09-02T08:57:34.000Z
**Category:** Research
**Topics:** agent, llm, benchmark, inference

## Summary

Ming Wu and Pengyuan Zhu's Agent Zero Memory architecture improves long\-term recall for LLM agents by separating memory into three concurrent search systems\. The framework runs an episodic events timeline, an entity\-event knowledge graph, and a curated documentary memory of durable facts to reach 95\.60% on LongMemEval and 93\.60% on LoCoMo\. Evaluation of eight backbone models found a 3\.4 point accuracy increase and a 30x reduction in per\-query costs\. Retrieval starts with an intent gate and a source router to narrow search buckets\. The "citation lock" mechanism ensures replies only cite evidence that a reader has actually opened, while indexed raw sources prevent data loss during bad extractions\. The authors have not yet performed ablation studies to determine if the accuracy gains result from the memory stores or the intent gate\.

## Sources

- [Story source](<https://huggingnews.com/ai/agent-zero-memory-cuts-llm-query-costs-30x-new-benchmark-high-of-9560per-50c0feb9>)
- [Story source](<https://x.com/dair_ai/status/2094953486047977860>)
- [Supporting source](<https://x.com/mrdj1968/status/2094956976161878121>)
- [Supporting source](<https://x.com/HEI/status/2094767863844237527>)

