Gated DeltaNet-2: splitting erase and write gates for linear attention

A new NVIDIA linear-attention model improves memory editing by separating channel-wise erase and write operations instead of sharing one scalar gate

New linear attention SoTA? Gated DeltaNet-2 from NVIDIA beats KDA and Mamba-3. Prior DeltaNet/KDA models used one scalar gate for both erasing old memory and writing new memory. This paper splits that into channel-wise erase and write gat
Ranked #16 on backlist 2026-05-23 (23 May 2026 UTC) · by (alphaXiv) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.