Paper Detail

C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks

Xueshu Chen, Yan Wang, Zihao Xue, Jiefu Li, Zhenfang Liu, Jayden Chen, Zhen Bi, Jungang Lou

arxiv Score 15.2

Published 2026-09-24 · First seen 2026-09-26

General AI

Abstract

Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records. At query time, budgeted routing selects useful index pages and expands their associated source evidence under a fixed reader budget. Together, these mechanisms establish a compact, provenance-preserving multimodal memory organization for cross-session long-horizon tasks, retaining temporal distinctions and source links required for reliable downstream reasoning. Code is available at https://github.com/HuzhouNLP/C3M.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{chen2026c3m,
  title = {C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks},
  author = {Xueshu Chen and Yan Wang and Zihao Xue and Jiefu Li and Zhenfang Liu and Jayden Chen and Zhen Bi and Jungang Lou},
  year = {2026},
  abstract = {Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that maintains a bounded active index over persistent source text-image evidence. Relation-aware updates consolidate safe redundancy while preserving complementary and incompatible records.},
  url = {https://arxiv.org/abs/2609.29735},
  keywords = {cs.AI, cs.CL, cs.CV, cs.IR},
  eprint = {2609.29735},
  archiveprefix = {arXiv},
}

Metadata

{}