Mandol in Practice: Building Agent Long-Term Memory with Zero LLM Calls

·Toolin Editorial Team

The Institute of Software, Chinese Academy of Sciences, open-sources Mandol, which unifies KV/vector/graph storage via SemanticMap + SemanticGraph — zero LLM calls at retrieval, 5.4x faster retrieval, and best-in-class results on both LoCoMo and LongMemEval.

Mandol in Practice: Building Agent Long-Term Memory with Zero LLM Calls

Memory for long-running conversational agents is a stubborn old problem: traditional setups keep vector stores and graph databases separate, with high cross-store I/O latency; RAG retrieval introduces noise, misses related clues, and offers no control over the token budget. Mandol, open-sourced by the Institute of Software at the Chinese Academy of Sciences together with Microsoft Research, proposes an "agglomerative" approach — unifying fragmented storage and representation into one memory-native architecture, with a retrieval process that never calls an LLM, plus the best overall accuracy on two mainstream long-term conversation benchmarks.

What Mandol Solves

Two core pain points of existing agent memory systems:

  1. Fragmented storage: the heterogeneous combination of a vector store plus a graph database scatters memory information, and cross-store I/O drives up latency
  2. Poor retrieval quality: common RAG setups introduce noise, miss related clues, and lack token budget control, dragging down LLM accuracy and efficiency

Mandol's answer is to agglomerate all memory representation and storage into a single unified architecture.

Three Core Components

1. Hierarchical Memory Model

Memory is split into two tiers, uniformly represented as a structured semantic graph:

  • Basic layer: holds the raw memory information
  • Abstract layer: agglomerates basic memories into traceable abstract memories

This layering lets the agent both look up raw facts and reason over higher-level abstract memories, instead of only fishing for fragments via vector similarity.

2. Agglomerative Semantic Data Structures

This is Mandol's engineering core: a SemanticMap + SemanticGraph combination that natively fuses three structures:

  • Key-value
  • Vector
  • Graph

It also provides unified hybrid retrieval operators that eliminate cross-store I/O. Retrieval no longer has to bounce between a vector store and a graph store — every operation completes on the unified data structure.

3. Quantitative Query Mechanism

The entire query pipeline makes no LLM calls and includes three capabilities:

  • Query-adaptive routing
  • Quantitative denoising and conflict resolution
  • Token-constrained context generation

That means the retrieval stage burns no tokens, keeping both cost and latency under control.

Before You Start

  • Paper: https://arxiv.org/abs/2606.29778
  • Authors: Institute of Software, Chinese Academy of Sciences (Yuhan Zhang et al.) + Microsoft Research (Wentao Wu)
  • Subject classification: cs.DB / cs.AI / cs.CL / cs.IR

Measured Performance (Reference)

MetricMandol result
LoCoMo (long-term conversation benchmark)Best overall accuracy among representative systems
LongMemEval (long-term conversation benchmark)Best overall accuracy among representative systems
Retrieval speedup5.4x (at 10 QPS concurrency)
Insertion speedup4.8x (at 10 QPS concurrency)
Consumer hardware latencyStays low

The comparison targets are "representative agent memory systems" — in other words, among mainstream agent memory systems, Mandol posts the best or tied-best results on both the accuracy and speed dimensions.

How to Apply Mandol's Approach to Your Own Agent

Implementation details require the full arXiv paper, but if you want to borrow this approach for your own agent, you can restructure the memory module along these three layers.

Step 1: Unify the Storage Layer

Merge the previously separate vector store and graph store into one unified structure:

  • Each memory carries KV fields, a vector embedding, and graph node/edge information at the same time
  • Don't store the same memory twice in two stores and join later — that's exactly where the traditional approach's latency comes from
  • Replace it with a SemanticMap (fast lookup by key) + SemanticGraph (traverse associations) combination

Step 2: Layered Abstraction

Don't store only raw memories; periodically aggregate basic memories into abstract ones:

  • Basic tier: raw facts, user statements, and events from each conversation
  • Abstract tier: preferences, habits, and long-term relationships generalized from multiple basic memories
  • Abstract memories must be traceable — you should be able to trace back which basic memories they were aggregated from, making correction and updating easy

Step 3: Zero-LLM Retrieval

Make retrieval a pure computational pipeline:

  • Routing: choose KV lookup, vector similarity, graph traversal, or a mix depending on query type
  • Denoising: filter out irrelevant or conflicting memories with quantitative methods (rather than asking an LLM to judge)
  • Token control: a hard token budget, truncated by relevance score, to avoid stuffing too much context into the main LLM

Together, these three steps are the core of the Mandol paper's "retrieval without involving LLMs."

When It Fits and When It Doesn't

A good fit for the Mandol approach:

  • Long-running conversational agents (companions, support, personal assistants) that need cross-session memory
  • Memory volumes large enough that the vector store + graph store combo starts dragging on latency
  • Sensitivity to per-query token cost, with no desire to burn LLM calls during retrieval either

Not a great fit:

  • Single-turn QA and memory-free scenarios
  • Tiny memory volumes (a few hundred entries at most) where simple vector search already suffices
  • Tasks that lean heavily on LLM reasoning for memory consolidation (Mandol's advantage is zero-LLM retrieval, not consolidation)

Verifying the Results

After adopting the Mandol approach, focus on three metrics:

  1. Retrieval latency: should improve by an order of magnitude versus the old "vector store + graph store join" setup (5.4x in the paper)
  2. Insertion throughput: writing new memories shouldn't become the bottleneck (4.8x speedup in the paper)
  3. Long-term conversation accuracy: should beat your baseline on standard sets like LoCoMo / LongMemEval

If your agent's latency stays high under concurrency, check whether cross-store I/O was truly eliminated; if accuracy doesn't improve, check whether the layered abstraction is genuinely agglomerating rather than just storing one more redundant copy.

Primary sources: