FlashMemory: 90% smaller KV cache at 500K context

Lookahead Sparse Attention claims long-context memory compression without the usual accuracy collapse, attacking one of the main inference cost centers

FlashMemory-DeepSeek-V4 Lookahead Sparse Attention cuts the KV cache by over 90% at 500K context, compressing it to just 13.5% of full size while maintaining or improving accuracy on RULER, LongBench-v2, and LongMemEval.
Ranked #9 on backlist 2026-06-14 (14 Jun 2026 UTC) · by (DailyPapers) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.