Paper Detail

Reward-Free Continual Adaptation for Resilient Space Robots

Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez

arxiv Score 10.3

Published 2026-08-24 · First seen 2026-08-26

Research Track A · General AI

Abstract

Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While continual reinforcement learning offers a promising mechanism for online adaptation, it inherently requires access to a reward signal during deployment. However, precise reward computation in space is often infeasible due to the lack of external tracking systems and the overall complexity of the environment. To address the challenge of unobservable rewards, we introduce a reward-free continual learning framework that leverages latent-state world models. By pre-training a model-based agent across diverse simulations, the world model learns a robust predictor of the reward structure within its latent space. Upon deployment to an environment with severe hardware degradation, we freeze the observation encoder and reward predictor to update only the transition dynamics of the world model through unsupervised rollouts. By training the policy entirely on imagined trajectories generated by this updated world model, the agent adapts to altered dynamics without receiving new rewards. We demonstrate our approach across simulated planetary traversal, orbital navigation, and precision assembly tasks subjected to severe morphological failures.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{orsula2026reward,
  title = {Reward-Free Continual Adaptation for Resilient Space Robots},
  author = {Andrej Orsula and Miguel Olivares-Mendez and Carol Martinez},
  year = {2026},
  abstract = {Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While continual reinforcement learning offers a promising mechanism for online adaptation, it inherently requires access to a reward signal during deployment. However, precise reward computation in space is often infeasible due to the lack of external tracking systems and the overall complexity of the environment. To address the challenge of unobservable rewards, we i},
  url = {https://arxiv.org/abs/2608.23452},
  keywords = {cs.RO, cs.AI, cs.LG},
  eprint = {2608.23452},
  archiveprefix = {arXiv},
}

Metadata

{}