Paper Detail

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

Jiajun Shi, Siyuan Tao, Yuhao Wu, Zexuan Wang, Jingyuan Zhang, Jiaheng Liu, Xinping Lei, Xinrong Zhang, Siyuan Fang, Zhewen Tan, Tianle Cai, Junhao Fang, Jiameng Huang, Yueyang Wang, Jinkai Liu, Yuxuan Zhang, Jian Yang, Zhoujun Li, Shen Yan, Wenhao Huang, Ge Zhang

arxiv Score 17.3

Published 2026-08-31 · First seen 2026-09-01

Research Track A · General AI

Abstract

Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use that experience to improve future decisions. We introduce \textbf{S\textsuperscript{3}Gym}, an interactive benchmark for evaluating LLM self-improvement through three coupled capabilities: \textbf{Self-Testing}, \textbf{Self-Judging}, and \textbf{Self-Improvement}. S$^3$Gym separates permissive exploration from strict held-out evaluation and instantiates this protocol in seven text-based games with executable environment verifiers. We evaluate three pathways for incorporating interaction experience: direct History ICL, score-conditioned Summary Memory, and parameter Training. Our experiments reveal that self-improvement is neither automatic nor uniform. Context-level experience improves performance for several model--game pairs, but the most effective pathway depends strongly on the task structure: summaries are beneficial when experience can be compressed into reusable strategic rules, yet often underperform raw history when success depends on precise, state-contingent information. Parameter training produces substantial gains on some tasks, but also exhibits unstable improvement and severe negative transfer on others. These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies. S$^3$Gym provides a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{shi2026s3gym,
  title = {S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?},
  author = {Jiajun Shi and Siyuan Tao and Yuhao Wu and Zexuan Wang and Jingyuan Zhang and Jiaheng Liu and Xinping Lei and Xinrong Zhang and Siyuan Fang and Zhewen Tan and Tianle Cai and Junhao Fang and Jiameng Huang and Yueyang Wang and Jinkai Liu and Yuxuan Zhang and Jian Yang and Zhoujun Li and Shen Yan and Wenhao Huang and Ge Zhang},
  year = {2026},
  abstract = {Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use that experience to improve future decisions. We introduce \textbackslash{}textbf\{S\textbackslash{}textsuperscript\{3\}Gym\}, an interactive benchmark for evaluating LLM self-improvement through three coupled capabili},
  url = {https://arxiv.org/abs/2608.31100},
  keywords = {cs.CL, large language models, self-testing, self-judging, self-improvement, in-context learning, summary memory, parameter training, negative transfer, executable environment verifiers, huggingface daily},
  eprint = {2608.31100},
  archiveprefix = {arXiv},
}

Metadata

{}