Paper Detail

Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers

Zonglin Yang, Ziming Zhao, Wei Tang, Xunyu Jiang, Yihong Liu, Tailin Chen, Zifu Yu, Jiayu Liu

arxiv Score 9.9

Published 2026-09-09 · First seen 2026-09-10

Research Track A · General AI

Abstract

Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active ($0.772 \pm 0.020$) but collapse at zero gate ($0.095 \pm 0.009$). Smooth fade-to-zero training preserves high zero-gate accuracy ($0.734 \pm 0.028$), whereas forced-zero training, hard switching, and post hoc continuation fail to recover the same effect. The pattern also appears on Markov induction. Linear regression ICL provides a boundary case because zero-gate training can learn that task directly. Mechanistic traces show that circuit consolidation occurs after the gate reaches zero, even though the responsible heads vary across seeds. These results suggest that circuit removability in small discrete retrieval tasks depends on the training trajectory, not just the final architecture.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{yang2026training,
  title = {Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers},
  author = {Zonglin Yang and Ziming Zhao and Wei Tang and Xunyu Jiang and Yihong Liu and Tailin Chen and Zifu Yu and Jiayu Liu},
  year = {2026},
  abstract = {Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active (\$0.772 \textbackslash{}pm 0.020\$) but collapse at zero gate (\$0.095 \textbackslash{}pm 0.009\$). Smooth fade-to-zero training preserves high z},
  url = {https://arxiv.org/abs/2609.10287},
  keywords = {cs.LG, cs.NE},
  eprint = {2609.10287},
  archiveprefix = {arXiv},
}

Metadata

{}