15 neuroscience-inspired mechanisms across 7 memory layers. Open-access preprint on arXiv. All documented.
“Memory is not storage — it is a living process of forgetting, consolidation, and rediscovery. We translated this process into software.”
Our technical paper is publicly available as an open-access preprint on Zenodo and arXiv, with a defensive disclosure on Elsevier TDCommons.
We present ZenBrain, a neuroscience-inspired 7-layer memory architecture integrating 15 mechanisms grounded in peer-reviewed neuroscience: 9 foundational algorithms plus a Predictive Memory Architecture (PMA) with NeuromodulatorEngine, ReconsolidationEngine, TripleCopyMemory, PriorityMap, StabilityProtector, and MetacognitiveMonitor. Evaluated across ten experiments on LoCoMo, MemoryAgentBench, MemoryArena, and the LongMemEval-S Full-500 replication. On LongMemEval-500, three of nine head-to-head judge comparisons hold for ZenBrain, all three against A-Mem (Bonferroni-corrected p ≤ 6.2e-31); the remaining six, against Letta and Mem0, are ties at the paper's own criterion, judged at equal evidence depth and under version-matched judges. No comparison is lost. Under the official binary judge, ZenBrain reaches 91.3 % of long-context-oracle accuracy at 1/109.6 of the per-query token budget — the oracle beats ZenBrain by only 4.5 pp while using ~109.6× more tokens and no memory architecture. Sleep consolidation: +37 % stability, −47.4 % storage. TripleCopyMemory retains 91.2 % strength at 30 days. The full 15-mechanism ablation reveals a cooperative survival network where 9 of 15 mechanisms become individually critical under stress (decay=0.25, 60 days). All experiments are reproducible with seeded PRNG; the implementation is released as open-source packages under the @zensation npm scope.
Eight technologies shipped together in production in this combination.
Idle-time memory consolidation in production — inspired by hippocampal replay (Stickgold & Walker 2013). Weak connections are pruned, stable ones strengthened.
Among the systems compared here — Mem0, Letta, Zep — none ships sleep-time consolidation; the published designs for it remain proposals.
Source: zenbrain (opens in new tab)Unified orchestrator for 7 memory layers in production — from working memory to cross-context memory. Based on Global Workspace Theory (Baars 1988).
Mem0 ships 2 layers, Letta 3, Zep 2. ZenBrain ships 7 — among the deepest memory architectures in open source today.
Source: zenbrain (opens in new tab)Meta-agent that creates retrieval plans before any search is executed. Heuristic-first with LLM fallback, max 4 dependent steps.
A dedicated planning layer ahead of retrieval, with a heuristic gate: simple queries never reach the planner and cost no LLM call.
Source: ZenAI snapshot (8 May 2026) (opens in new tab)Structured 3-round debate protocol when agents disagree. Challenge → Response → Resolution with automatic escalation.
After three rounds without consensus the protocol escalates to a human instead of computing a majority.
Source: ZenAI snapshot (8 May 2026) (opens in new tab)Automatic knowledge gap detection with quantified gap score. Analyzes query history, fact density, and confidence — recommends targeted actions.
Gap detection runs on its own and proposes actions rather than waiting for a question. Inspired by Loewenstein's information gap theory (1994).
Source: ZenAI snapshot (8 May 2026) (opens in new tab)User intent prediction from temporal and sequential patterns. Learns from prediction errors — the more often wrong, the better the next prediction.
Among the systems compared here — Mem0, Letta, Zep — none predicts user intent from behavioral patterns within the memory layer itself.
Source: ZenAI snapshot (8 May 2026) (opens in new tab)3-level recursive self-improvement with formal safety bounds. Level 0 optimizes knowledge, Level 1 optimizes Level-0 strategies, Level 2 optimizes Level-1 parameters.
3-level recursion with immutable core properties, daily budgets, and automatic rollback on quality regression: self-improvement runs only inside those bounds.
Source: ZenAI snapshot (8 May 2026) (opens in new tab)Detection and merging of entities across 4 isolated contexts (Operations, Finance, People, Strategy). Bayesian confidence updates on conflicts.
Among the systems compared here — Mem0, Letta, Zep — none manages entity identity across isolated contexts.
Source: zenbrain (opens in new tab)An accessible, interactive depiction of the memory-based system — from the 7-layer memory and its neuroscience inspiration to how the agents work together. With a toggle between plain-language and technical explanations.
The idea in one sentence: collective intelligence on a cognitive architecture — humans and machines together, on a structure that keeps knowledge. The explorer below makes exactly that architecture tangible.
These mechanisms run for real — run the open library live →
Interactive depiction · best viewed on desktop
A working system sits on a narrow edge. Change the survival rule step by step and watch the grid collapse, carry structure, or clog — live. Most rules fail; only a narrow band holds.
Why quality is not optional for us: the very efficiency that makes a system valuable amplifies every rule you give it — for better and for worse.
The full story — why quality is not optional →
Interactive · click the neighbour counts on or off · best on desktop
The HiMeS architecture, inspired by the Atkinson-Shiffrin model (1968) and modern cognitive science.
Active focus — 7±2 items per Miller's Magical Number. Fastest access, shortest lifespan.
2026Session context and conversation continuity. Survives the current session.
2026Concrete experiences with emotional tagging. 400+ keyword lexicon (DE+EN) for arousal/valence scoring.
2026Factual knowledge with FSRS scheduling. Spaced repetition optimizes recall timing — 30% better than SM-2.
2026Workflows and skills. Tool chains are analyzed and optimized.
2026Immutable foundations following the Letta pattern. Pinned facts that are never forgotten.
2026Knowledge and entities linked across isolated contexts — with Bayesian confidence updates on conflict.
2026Optimal review timing at ~90% target retention
Exponential decay with configurable half-life
Co-activated facts strengthen connections (×1.09/activation)
Prevents runaway growth of edge weights
Hippocampal replay with +50% stability boost
Weak connections pruned during sleep
Emotional memories decay 2.7× slower
Confidence updates across the entire knowledge graph
Conscious access through competitive context assembly
Entropy-based prioritization of new facts
Systematic detection of missing knowledge
7±2 active items in working memory
Dopamine, NE, 5-HT, ACh — four channels with tonic + phasic dynamics
Memory destabilizes on retrieval — four PE-gated update modes with rollback
Three traces with divergent dynamics: fast (4h), medium (14d), deep (logarithmic)
6 strategies, dynamically selected per query. Not one pipeline — an adaptive system.
↻ Self-RAG Critique: loops back if confidence < 0.5
Meta-agent plans retrieval steps before execution. Heuristic-first, LLM fallback.
Event subgraph + semantic graph + community summaries. 5 parallel strategies.
Automatic reformulation at confidence < 0.5. 4-component scoring.
Hypothetical answer → embedding → search. Auto-detection with 5s timeout.
Chunk enrichment per Anthropic method. +67% retrieval accuracy.
BullMQ worker monitors drift >10%. Automatic cache invalidation.
Multi-agent orchestration with structured debate, dynamic team composition, and recursive self-improvement.
3 rounds: challenge → response → resolve
3-round debate on disagreement. Challenge → Response → Resolution.
5 specialist agents, automatically composed by task type.
Recursive self-improvement with daily budgets, sandbox tests, and auto-rollback.
Pause/resume with state checkpointing. Long-running tasks over days.
Agent-to-agent communication per Google standard. /.well-known/agent.json discovery.
Automatic behavior detection without explicit labeling.
Curiosity, prediction, metacognition — three pillars of cognitive intelligence.
Quantified gap score: query frequency × fact density × confidence × RAG quality.
Temporal + sequential pattern recognition. Learns from prediction errors.
Confidence calibration, confusion detection, capability profiling.
4-tier thinking budgets: 1K→16K→64K→128K tokens. Auto-detection. 60-80% cost savings.
On LongMemEval-500, ZenBrain reaches 91.3% of long-context-oracle accuracy at a per-query token budget of 1/109.6.
In our own measurements on LongMemEval-500, three of nine pairwise comparisons hold, all three against A-Mem, Bonferroni-corrected p ≤ 6.2 × 10⁻³¹ across three independent LLM judges. The six against Letta and Mem0 are ties at the paper's own criterion, judged at equal evidence depth and under version-matched judges; none is lost.
ZenBrain is open source. All algorithms, all tests, all documentation — openly available.
A vector store searches for similarity in a flat index. ZenBrain models memory as a process: what is used repeatedly consolidates, what is left alone fades. Seven layers separate working memory, episodic, semantic and procedural knowledge, and a consolidation phase at rest condenses experience into knowledge. The difference shows where context has to hold for weeks rather than for one conversation.
Because different kinds of memory need different rules. Working memory has to be fast and volatile, semantic knowledge durable and condensed, procedural memory retrievable without deliberate search. Each layer carries its own decay and consolidation logic; fifteen mechanisms in total, nine foundational plus six in a predictive architecture.
Yes, and each one with its source in the paper: FSRS for spacing intervals, Hebbian learning, the Ebbinghaus forgetting curve, Bayesian confidence propagation, a two-factor synaptic model, a vmPFC-coupled prediction-error update, and sleep consolidation. Grounded here means in peer-reviewed neuroscience, not in an analogy.
On LongMemEval-500, at the same token budget and judged by three independent LLM judges, three of nine head-to-head answer-quality comparisons hold for ZenBrain, all three against A-Mem. The remaining six, against Letta and Mem0, are ties at the paper's own criterion, with none lost. All nine contrasts are Bonferroni-corrected and judged under version-matched judges. It reaches 91.3 percent of the accuracy of a full-context oracle at 1/109.6 of the per-query token budget.
Yes, and that is the point. The memory core is open under Apache 2.0, the publications carry DOIs on Zenodo, and the benchmark page gives every figure its method, source and verification path. Anyone who doubts a number should be able to recompute it without asking us.
The preprint itself (arXiv:2604.23878) is not peer-reviewed. The neuroscience it builds on is. We state that plainly, because the distinction matters — and because open evidence has to carry a claim, not a reference to a process.
More from this research
Core and application fields, on a shared ethics foundation.