DeepSeek Engram adds a second sparse axis to large language models. Mixture-of-Experts chooses which computation to activate; Engram chooses which static memory rows to retrieve. The important shift is infrastructural: capacity can move from GPU-resident computation toward deterministic lookups that may be prefetched from host memory. Whether that becomes a real deployment gain depends on bandwidth, overlap, locality, and evidence beyond the authors' experiments.
Reading time: 10 minutes · About 1,900 words
TL;DR
- MoE provides conditional computation. Engram provides conditional memory through token-derived N-gram lookups.
- The paper's compute-matched experiments found a U-shaped allocation curve: assigning about 20% to 25% of sparse capacity to Engram performed best in the two tested regimes.
- Engram-27B beat an iso-parameter, iso-activated-parameter MoE-27B across many reported tasks, but these remain author-run pretraining results.
- A 100B-parameter Engram table offloaded to host DRAM reduced throughput by at most 2.8% in the paper's narrow H800 prototype benchmark. This is evidence for feasible overlap, not a universal production guarantee.
- DeepSeek V4.1-Flash includes a 196B-parameter Engram component, while the public Engram repository provides a demonstration that mocks core backbone components rather than a full reproduction.
Conditional memory is different from conditional computation
An MoE layer routes each token through a small subset of experts. Total parameter capacity grows while activated compute per token stays bounded. Yet the model still spends neural computation reconstructing common names, phrases, and other static local patterns.
Engram introduces a separate primitive. It converts recent token N-grams into deterministic hash addresses, retrieves a small number of learned embedding rows, and uses a context-aware gate to fuse the retrieved value into the model's hidden state. The lookup count stays constant as the table grows, so the authors describe it as an O(1) memory operation.
The distinction matters:
| Sparse axis | Address depends on | Main resource | Best suited to |
|---|---|---|---|
| MoE conditional computation | Runtime hidden state and router | Accelerator compute and expert communication | Dynamic transformation and reasoning |
| Engram conditional memory | Input token IDs and static hashing | Memory capacity and data movement | Recurrent static token patterns |
Memory does not replace computation. The paper's allocation experiment found the worst direction at either extreme: pure MoE lacks a dedicated lookup path, while an Engram-heavy model loses dynamic expert capacity.
What the allocation experiments found
DeepSeek tested two smaller compute regimes while keeping total-to-activated sparsity near ten. The only changed allocation was how inactive parameter capacity was split between routed experts and Engram slots.
Both experiments produced a U-shaped loss curve. The reported optimum assigned roughly 75% to 80% of sparse capacity to MoE, leaving about 20% to 25% for Engram. In the approximately 10B-parameter regime, validation loss moved from 1.7248 for pure MoE to 1.7109 near the optimum.
The larger pretraining comparison froze 262B training tokens and 3.8B activated parameters. MoE-27B and Engram-27B each had 26.7B total parameters; the Engram variant replaced some routed experts with a 5.7B-parameter memory table. The paper reports Engram-27B gains including:
- MMLU: 57.4 to 60.4;
- CMMLU: 57.9 to 61.9;
- BBH: 50.9 to 55.9;
- ARC-Challenge: 70.1 to 73.8;
- HumanEval: 37.8 to 40.8;
- MATH: 28.3 to 30.7.
Those comparisons support the paper's central claim under its training setup: some sparse capacity was more valuable as explicit memory than as additional routed experts. They do not yet establish the same allocation ratio for different tokenizers, datasets, model scales, or hardware.
What actually moves out of the GPU
Engram's addresses are known once the token sequence is known. During inference, the system can compute upcoming memory indices before the Engram layer executes, fetch rows from host DRAM over PCIe, and overlap that transfer with earlier Transformer blocks.
This changes the resource map:
token IDs
→ deterministic N-gram addresses
→ host-memory embedding rows
→ asynchronous prefetch over PCIe
→ context gate on GPU
→ fused residual state
The table capacity can reside outside HBM, but the selected rows still move to the accelerator and the gate still runs in the model. Training has a different cost: the paper shards tables across GPUs and uses all-to-all communication for active rows and gradients.
The architecture also creates a placement trade. Early insertion can relieve the backbone from reconstructing static patterns before it spends many layers doing so. Deeper insertion provides more preceding compute time to hide host-memory latency and better contextual states for gating. The paper's 12-layer ablation favored an early Layer 2 insertion, showing that the modeling optimum and the systems optimum must be solved together.
The 2.8% overhead claim has a narrow scope
The paper tested a 100B-parameter Engram table fully stored in host DRAM. Its prototype used a nano-vLLM-based harness, one NVIDIA H800, 512 sequences, and sequence lengths uniformly sampled from 100 to 1,024 tokens. The test used dense 4B and 8B backbones to avoid MoE communication as a confounder.
Reported throughput changed as follows:
| Backbone | Baseline | With 100B Engram in host DRAM | Difference |
|---|---|---|---|
| Dense 4B | 9,031.62 tok/s | 8,858.28 tok/s | about 1.9% lower |
| Dense 8B | 6,315.52 tok/s | 6,140.02 tok/s | about 2.8% lower |
This is credible evidence that deterministic prefetch can hide most transfer time in that setup. It does not prove negligible overhead for low batch sizes, longer sequences, NUMA effects, saturated PCIe links, distributed serving, tail latency, or concurrent models sharing host bandwidth. The benchmark measures aggregate throughput, while interactive serving often fails first at P95 or P99 latency.
The paper also describes a hierarchy that caches frequent N-grams in faster tiers and leaves the long tail in slower storage. Its throughput experiment deliberately forced all retrieval through PCIe, so it is a conservative test of one link. It still leaves the cache policy and production locality distribution to be validated.
Evidence levels: paper, code, and deployed model
Three public artifacts answer different questions.
First, the paper provides the architecture, controlled training comparisons, ablations, and prototype systems benchmark. These are author-run results.
Second, the official GitHub repository provides a standalone demo of Engram's data flow. Its README explicitly says the implementation mocks Attention, MoE, mHC, and other standard components. It is useful for understanding lookup and gating, but it cannot reproduce the training table or host-offload benchmark by itself.
Third, DeepSeek's V4.1-Flash model card says the released model contains 196B Engram parameters, sparsely accessed through token-based lookup. This establishes adoption in a released architecture. It does not isolate how much of V4.1's end-task performance comes from Engram because V4.1 also changes attention, KV compression, residual mixing, speculative decoding, data, and post-training.
The adoption timeline needs one correction to early speculation: the original DeepSeek V4 model page did not list Engram; V4.1-Flash is the release that explicitly does. Qwen3.8-Flash-Next then provides a second industrial signal. Its official materials describe a 125B backbone plus 51B of N-gram embeddings with 6B parameters activated per token. Qwen cites Engram and uses related multi-head hashing and contextual gating, but calls the component N-gram Embedding. It is an Engram-influenced conditional-memory design rather than evidence of an identical transplant.
Public evidence therefore reaches author experiments, a demonstration artifact, and production-model adoption. An independent end-to-end reproduction of the paper's allocation law and offload economics remains a separate evidence level.
Early external results also warn against treating lower language-model loss as guaranteed downstream gain. A small-model replication on Qwen-family models reports 22% to 30% lower held-out perplexity across four experiments but no consistent improvement on HellaSwag, PIQA, or ARC-Challenge. This is a single community replication on 0.5B to 4B models, so it does not invalidate DeepSeek's large-scale result. It does show that scale and evaluation target matter, and that perplexity alone is an insufficient deployment gate. Separate 2026 research on tokenizer-agnostic Engram and Memory Grafting further modifies the lookup design, which is evidence that conditional memory is an active design space rather than a settled recipe.
Engram is not a replacement for RAG
Engram stores learned parametric rows tied to token patterns. Its contents are established during training and addressed from the token stream. Retrieval-augmented generation searches an external corpus that can be updated, cited, access-controlled, and inspected independently of model weights.
Use Engram's conceptual role for frequent static associations and model capacity. Use RAG or tools when freshness, provenance, deletion, permissions, or document-level evidence matters. Combining them is plausible: Engram can reduce repeated internal reconstruction while external retrieval supplies current and attributable knowledge.
A deployment benchmark for the second sparse axis
Freeze model quality, hardware, and workload before evaluating memory offload. Compare at least:
| Dimension | Variants |
|---|---|
| Batch and concurrency | 1, normal load, saturation |
| Sequence shape | short prompts, long context, long decode |
| Memory locality | hot N-grams, mixed traffic, cold tail |
| Topology | local DRAM, NUMA remote memory, shared PCIe load |
| Task | factual recall, multilingual phrases, reasoning, code, long-context retrieval |
| Failure | delayed fetch, bandwidth contention, missing shard, worker restart |
Measure HBM and host-memory use, bytes transferred per token, prefetch hit rate, overlap time, accelerator stall time, throughput, time to first token, P50/P95/P99 latency, quality by task, training communication, and total cost per successful request.
The main decision is not whether a table fits in CPU memory. It is whether memory access remains hidden under the actual compute window. On a faster backbone or a congested link, the overlap budget shrinks. After HBM pressure falls, PCIe bandwidth, host capacity, cache scheduling, or training all-to-all can become the next bottleneck.
FAQ
What is DeepSeek Engram?
Engram is a conditional-memory module that hashes recent token N-grams into learned embedding tables, then gates retrieved values into the Transformer state.
Why is it called a second sparse axis?
MoE sparsifies which computation runs. Engram sparsifies which memory rows are read. The two capacities can be allocated independently under a fixed activated-compute budget.
Does Engram move all model memory to CPU RAM?
No. It allows the large static Engram table to reside in host memory during inference. The backbone, active computation, selected rows, and gating still use the accelerator.
Is the reported 2.8% overhead independently verified?
The figure comes from the authors' H800 prototype benchmark with dense backbones and high concurrency. It has a precise setup and should be reproduced on the target serving topology before production use.
Is Engram the same as RAG?
No. Engram is trained parametric memory addressed by token patterns. RAG retrieves external, updateable documents and can provide provenance.
References
- DeepSeek paper: Conditional Memory via Scalable Lookup
- DeepSeek official Engram repository
- DeepSeek-V4.1-Flash model card
- Qwen3.8-Flash-Next official repository
- Qwen3.8-Next technical report
- Small-model Engram replication and negative downstream result
- Tokenizer-Agnostic Engram Module
- Memory Grafting: Offline Conditional Memory
- Related: DeepSeek V4.1 SWA bounded replay audit
- Related: Latent workspace versus external memory
Engram's strongest idea is the separation of resources. Static recall, dynamic reasoning, accelerator compute, and memory capacity no longer need to scale as one block. Its strongest open question is equally concrete: how reliably can a real system keep that new memory path off the critical latency path?