Paper Detail
Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Yuxuan Zhang, Xinping Lei, Junting Zhou, Zexuan Wang, Yuchen Wu, Huan Zhou, Duo Wang, Yinzhu Piao, Yongchang Peng, Yunfeng Shi, Jin Chen, Zuo Wang, Jinkai Liu, Jiaheng Liu, Wenxuan Zhang, Shen Yan, Wenhao Huang, Ge Zhang
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.
No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.
No ranking explanation is available yet.
No tags.
@article{wu2026aspire,
title = {Aspire: Can Models Self-Evolve from Vague Goals?},
author = {Yuhao Wu and Jingyuan Zhang and Jiajun Shi and Yuxuan Zhang and Xinping Lei and Junting Zhou and Zexuan Wang and Yuchen Wu and Huan Zhou and Duo Wang and Yinzhu Piao and Yongchang Peng and Yunfeng Shi and Jin Chen and Zuo Wang and Jinkai Liu and Jiaheng Liu and Wenxuan Zhang and Shen Yan and Wenhao Huang and Ge Zhang},
year = {2026},
abstract = {Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPI},
url = {https://arxiv.org/abs/2608.31111},
keywords = {cs.CL, LLM self-evolution, vague-goal-driven self-evolution, ASPIRE, model-weight evolution, agent-harness evolution, hidden evaluation, self-evaluation, huggingface daily},
eprint = {2608.31111},
archiveprefix = {arXiv},
}
{}