Paper Detail

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel

huggingface Score 17.5

Published 2026-08-31 · First seen 2026-09-03

General AI

Abstract

We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{moriya2026foldingagent,
  title = {FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos},
  author = {Maya Moriya and Sigal Raab and Yael Vinker and Tali Dekel},
  year = {2026},
  abstract = {We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parame},
  url = {https://huggingface.co/papers/2609.00377},
  keywords = {Vision-Language Model, parametric folding programs, origami demonstration videos, geometric transitions, physical plausibility, cross-modal retrieval, parametric folding actions, crease patterns, multi-step folding, PurelandFold, huggingface daily},
  eprint = {2609.00377},
  archiveprefix = {arXiv},
}

Metadata

{}