FlashMemory: 90% smaller KV cache at 500K context
Lookahead Sparse Attention claims long-context memory compression without the usual accuracy collapse, attacking one of the main inference cost centers
FlashMemory-DeepSeek-V4 Lookahead Sparse Attention cuts the KV cache by over 90% at 500K context, compressing it to just 13.5% of full size while maintaining or improving accuracy on RULER, LongBench-v2, and LongMemEval.