Paper Detail

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

Tommaso Cerruti, Tim Rieder, George Rowlands, Lingfeng Jin, Imanol Schlag

huggingface Score 9.5

Published 2026-07-08 · First seen 2026-07-10

General AI

Abstract

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity. Our experiments center on 350M-parameter models trained for 15B tokens, and include optimizer and learning-rate comparisons, hybrid-versus-pure stack comparisons, sequence-length runtime measurements, larger DeltaNet runs at 1.3B and 3B parameters, and a small set of downstream evaluations. The reported speed results measure training throughput and iteration time; we do not provide an empirical inference-speed benchmark. Within the reported 350M-parameter, 15B-token sweep, Kimi Delta Attention with Muon reaches the lowest final validation loss, a pure Gated DeltaNet stack trained with AdamW has the highest normalized training throughput, hybrid stacks generally improve loss at a throughput cost, and Muon consistently lowers final validation loss relative to AdamW in the matched architecture settings we evaluate. We introduce and evaluate lightweight cross-layer routing mechanisms for DeltaNet-style memories. The most natural DeltaNet-inspired formulation, forwarding a lower layer's delta-rule write error into the next layer's value target, does not improve over matched baselines. Routing into the aligned hidden stream and forwarding the write value instead yields a modest improvement in the matched runs we report: Cross-Layer Value Routing (CLVR) lowers final validation loss for both DeltaNet and Gated DeltaNet.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{cerruti2026linear,
  title = {Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing},
  author = {Tommaso Cerruti and Tim Rieder and George Rowlands and Lingfeng Jin and Imanol Schlag},
  year = {2026},
  abstract = {Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write c},
  url = {https://huggingface.co/papers/2607.07953},
  keywords = {self-attention, softmax attention, recurrent linear-attention, DeltaNet, Gated DeltaNet, Kimi Delta Attention, Gated DeltaNet-2, recurrent-memory notation, memory decay, erase and write control, training throughput, implementation complexity, attention mechanism, sequence length, validation loss, optimizer, learning rate, hybrid stacks, pure stack, cross-layer routing, delta-rule write error, hidden stream, Cross-Layer Value Routing, CLVR, code available, huggingface daily},
  eprint = {2607.07953},
  archiveprefix = {arXiv},
}

Metadata

{}