Gated DeltaNet-2: splitting erase and write gates for linear attention
A new NVIDIA linear-attention model improves memory editing by separating channel-wise erase and write operations instead of sharing one scalar gate
New linear attention SoTA? Gated DeltaNet-2 from NVIDIA beats KDA and Mamba-3. Prior DeltaNet/KDA models used one scalar gate for both erasing old memory and writing new memory. This paper splits that into channel-wise erase and write gat