Treatmybrand


a Kainjoo SA Venture
Ch. du Vernay 14a
1196 Gland
+41.21.561.34.96
[email protected]

Support


Monday to Friday
8AM to 8PM
[email protected]
Back

DeepSeek’s Conditional Memory Innovates GPU Efficiency by Reducing Static Lookups in LLMs

Enterprises using large language models (LLMs) frequently waste costly GPU resources to retrieve static information like product names or contract clauses, leading to inflated infrastructure expenses. DeepSeek’s new research introduces “conditional memory” via the Engram module, which separates static pattern retrieval from dynamic reasoning. This reduces redundant GPU computation by handling static data through efficient hash lookups filtered by contextual gating.

DeepSeek found the optimal mix to be around 75% model capacity dedicated to dynamic reasoning and 25% for static data retrieval. This approach boosts performance on complex reasoning tests from 70% to 74% accuracy and knowledge retrieval from 57% to 61%.

Unlike agentic memory systems focused on episodic recall, conditional memory enhances internal model efficiency by providing native knowledge lookup capability, avoiding prolonged computations for basics like named entities. Its deterministic lookup mechanism allows embedding tables exceeding 100 billion parameters to be offloaded to CPU memory, dramatically easing GPU memory demands without significant slowdowns.

For enterprises, this suggests AI infrastructure can be rethought towards hybrid architectures balancing computation and memory, potentially shifting investment from expensive GPU scaling to memory-rich setups. The breakthrough indicates future AI models will achieve better reasoning performance with lower operational costs.

DeepSeek’s work charts a promising route for AI developers facing the dual challenge of escalating GPU costs and demand for smarter, more efficient LLMs.

Venturebeat
Venturebeat