AgentDreamerAgentic 3D Motion Reasoning for Text-Controllable Physics Simulation

Mijeong Kim1, Suho Park2, Junha Park2, Bohyung Han1,2

1ECE & 2IPAI, Seoul National University

Under review

Abstract

We propose AgentDreamer, an agent-guided optimization framework that aligns text instructions with simulated motion through explicit 3D motion planning. Given a text instruction and scene geometry, an off-the-shelf agent translates the instruction into a structured semantic motion prior, specifying what should move, in which direction, and in what temporal order. The agent represents this prior using parametric functions over a series of temporal chunks, yielding dense 3D trajectories. We then optimize the parameters of a differentiable MPM simulator such that its rollout best matches the prior while remaining constrained by the scene dynamics.

We further introduce Motion-IG, an information-theoretic metric for text–motion alignment based on masked text recovery. Experiments show that AgentDreamer improves text alignment over video-based guidance, producing simulations that more closely follow the requested physical behavior.

Animated comparison of a baseline and AgentDreamer for a flower rotating and bowing toward its pot.
Example: a flower rotates and bows toward its pot. AgentDreamer better follows the requested behavior in the shown comparison.

The key idea: agent-generated 3D motion

A visually plausible video reference can still miss the requested motion, such as sliding instead of rolling. AgentDreamer plans how the object and its parts should move before optimizing the simulation.

A text instruction and 3D scene are given to an agent, which reasons about the motion and produces 3D trajectories for a flower and soccer ball.
The agent directly constructs the 3D trajectory prior from text and scene geometry.
  • Scene-aware inputAnnotated views link sampled 3D control points to visible object parts.
  • Parametric motion planThe agent describes motion with functions over temporal chunks, yielding dense trajectories.
  • Physics optimizationThe trajectories supervise a differentiable MPM simulator while scene dynamics constrain the result.

Method

Text and a 3D scene enter an agent-based motion planner. Its parametric 3D trajectories become the target for differentiable physics optimization.

AgentDreamer method overview showing semantic scene annotations, temporal chunks of parametric motion, generated 3D trajectories, and differentiable MPM simulation.
Overview of AgentDreamer. The agent constructs a parametric 3D motion plan; its trajectories guide the simulator toward the requested behavior.

Qualitative Results

The examples below compare AgentDreamer with video-guided physics simulation methods under the same text instructions.

Rows compare PhysFlow, MotionPhysics, and AgentDreamer on a bouncing playdoh ball and a rocking fox.
Qualitative comparison on a deformable bouncing ball and a rocking fox. These examples are from the paper.

Quantitative Results

On 20 instructions across eight scenes, AgentDreamer obtains the highest average Motion-IG score among the compared methods. In a blinded evaluation with 25 participants, it is ranked first for text–motion alignment in 73.6% of rankings.

MethodOpen-weight Motion-IG ↑API-based Motion-IG ↑Human first rank ↑
PhysFlow0.52−0.2214.0%
MotionPhysics0.52−0.6312.4%
AgentDreamer0.71+0.4873.6%

Motion-IG measures how much simulated motion improves recovery of masked motion words beyond the remaining text and scene context. Open-weight scores average four evaluators; API-based scores average five evaluators. Higher is better. See the paper for the full protocol.

BibTeX

@misc{agentdreamer2027,
  title  = {AgentDreamer: Agentic 3D Motion Reasoning for Text-Controllable Physics Simulation},
  author = {Kim, Mijeong and Park, Suho and Park, Junha and Han, Bohyung},
  year   = {2027},
  note   = {Under review}
}