Under review

Keep the Effect, Drop the Actor

Programmable Effect-to-Execution
World-Action Models

Junyi Hu  ·  Zhewen He  ·  Yi Fang†

New York University Abu Dhabi

† corresponding author  ·  jh10472@nyu.edu

Copying the demonstrator fails on a held-out task, after a push, on a UR5e and under a new grasp; programming the effect succeeds in all four.
Keep the effect, drop the actor. A demonstration records how one arm moved and what happened to the world. Copying the demonstrator (top) ties the task to that motion: a published imitator fails the held-out task, replayed actions miss the target after a 5 cm push and fail on a UR5e, and a replay can only grasp as recorded. PEWAM keeps only the effect, as a program of object keypoints, a terminal state and two contact points (bottom), and re-solves the execution from the live scene: the same episode succeeds, the push is absorbed, a UR5e runs it without retraining, and rotating the contact points turns the grasp.

Abstract

A robot demonstration records two things in the same frames: what happened to the objects, and how one particular arm made it happen. We condition on the first. A demonstration is compiled into an effect program: the 3D keypoint trajectories of the objects that moved, two points marking where each was held, and the configuration the scene ends in, with the demonstrator removed. PEWAM, a 71.5M-parameter world-action model, generates effect, robot execution, action and terminal state as four streams with independent flow-matching times, so clamping a program and sampling the execution turns inference into programming, re-solved closed loop from the live scene.

On held-out LIBERO-Goal tasks, one demonstration's program completes 40 of 90 episodes, where the same backbone given a goal image or language, and published demonstration-conditioned methods, complete at most 19; on three of Meta-World's held-out classes it exceeds the best published results, though not on the five-class mean. Because a program is a set of coordinates, a person can edit it: the placement follows a shifted terminal state and the grasp turns with rotated contact points. The same program runs on four robot arms without retraining, and after a push, re-solving completes 23 of 60 episodes where replaying the demonstration completes 6. On a Franka arm, fine-tuned on real demonstrations of other tasks, programs compiled from single human videos complete 36 of 40 trials, against 22 for the same backbone conditioned on the video's last frame as a goal image.

Video

The method, the simulated results, and a real Franka arm programmed by single human videos (3:46, with narration).

The idea

Showing a robot what to do is often easier than describing it. A demonstration, however, records two things in the same frames: what happened to the objects, and how one particular arm made it happen. Watch someone put a bowl on a stove and you learn where the bowl should end up, roughly how it travels and where it can be held; you do not learn to bend your elbow the way they bent theirs.

The change in the world is the task, and the demonstrator's motion is one of many ways to produce it. A demonstration is therefore compiled into an effect program: keypoints of each relevant object tracked over time, two contact points recording where the object was held, and the configuration the objects end in. Nothing in the program refers to the robot.

01

An actor-free task

The task is written as object keypoints, contact points and a terminal state, with no robot geometry in it.

02

Editable, and portable across bodies

Moving the terminal state moves the placement, rotating the contact points turns the grasp, and one program runs on four arms.

03

Progress is a caused effect

Progress through the program counts only effects the robot has caused, so a push becomes a new observation rather than progress.

Method

For each moved object, the effect is the world-frame trajectory of the eight corners and the centre of its oriented box, plus two contact points at the finger-side contact patches at the moment of grasp. Each contact point is a fixed affine combination of the object's keypoints, so it moves rigidly with the object; rotating the contact points asks for another grasp of the same effect. The terminal state is the last frame of the effect. The effect and terminal state contain no robot geometry; the execution and the action describe only the body.

One DiT (8 layers, 71.5M parameters) processes observation, effect, execution, action and terminal-state tokens. Each stream draws its own flow-matching time, and a stream given clean contributes no loss, so the set of clean streams defines a query and one set of weights answers every conditional: inverse (effect and terminal state given, execution and action sampled), goal, forward and inverse dynamics.

Executing a program is the inverse query. The window of the program that a monotone pointer selects is clamped, the execution is sampled in 20 Euler steps (348 ms on one RTX A6000), the first four actions run, and the loop repeats from the new observation. A push moves an object just as a grasp does, so progress past the first programmed motion counts only once the gripper has closed and the object has moved 1 cm with it; before that, the program is re-anchored to where the object is. On another arm, the joint angles are re-expressed as the training arm's joints at the same end-effector pose, and no weights change.

PEWAM end to end: compiling a demonstration into an effect program, the training data, one model with a flow-matching time per stream, and the closed loop.
PEWAM end to end. (1) Compiling a demonstration into an effect program; (2) the training data; (3) one model with a flow-matching time per stream; (4) the closed loop. Bottom: contact as points, and three operators of the same loop.

Results

One demonstration programs a held-out task

The three held-out LIBERO-Goal tasks contribute no robot data to training. Given one demonstration of such a task, PEWAM completes 40 of 90 episodes from the full program and 49 from its terminal state alone, while the same backbone and data given a goal image, language or no specification complete none. With a bottle, a box and a carton removed from LIBERO-Object training, programs carry over to the removed objects: 67 of 90 from the program and 89 from the terminal state. On shelf-place, sweep-into and bin-picking, three of Meta-World's held-out classes, the program exceeds every published result.

LIBERO, held-out taskssuccess (%)
UWM21.1
Instant Policy5.6
ATM0.0
Zero-WAM0.0
PEWAM44.4
Replay of the demonstration*96.7
PEWAM, demonstrator's execution*94.4
Meta-World, held-outshelfsweepbin
Meta-World baseline0.024.70.3
DP312.716.79.3
Mamba Policy10.09.79.7
FreqPolicy7.015.711.0
FlowPolicy0.06.79.0
DAMI10.728.013.7
PEWAM22.2100.052.2

Success rate (%). Left: the three held-out LIBERO-Goal tasks, 90 episodes; each published method reads the same demonstration through its own interface. *Reference rows carry the demonstrator's motion. Right: Meta-World test classes from ML10 (shelf-place, sweep-into) and ML45 (bin-picking); PEWAM is the mean of three trainings, each over 30 episodes programmed by another episode's demonstration; published policies are fine-tuned on ten demonstrations of the class.

Program and final frame for four held-out tasks.
One program per held-out task, executed closed loop. Left of each pair: the program on its demonstration with the robot removed (object track white, terminal state yellow, contact points red and blue). Right: the last frame of the execution, with the object's path in white and its length in steps above. The last pair is a LIBERO-Object episode with a held-out object.

The program is editable

Translating the terminal state moves the realised placement with it: over 30 episodes shifted by ±5 and ±10 cm the slope is 0.71. Rotating the two contact points about the object's vertical axis turns the grasp the arm takes, in the commanded direction at every rotation and by 0.68 of it on average; with edited contacts the model completes 22 to 24 of 30 at ±20° and ±45°.

Terminal edits, contact edits and the resulting grasps.
Editing the program (bowl task). (a) Terminal edits on a 3×3 grid of contact rotation and terminal shift: realised against commanded shift, filled where the task predicate fires. (b) Contact edits: realised grasp rotation against commanded rotation of the two contact points, median of episode-paired shifts with 95% bootstrap bands. Contacts stored as affine weights (red) are read with the wrong sign. (c) The grasp with contacts rotated by ∓45°, from above.

One effect, many bodies and executions

One program runs on the Panda it was trained with and on a KUKA IIWA, a Kinova Gen3 and a UR5e, completing 27, 24, 22 and 16 of 30 bowl episodes with no retraining. Three sampling seeds succeed in 27, 28 and 28 of 30 through grasps a median 20.1° apart, and at least one seed succeeds in all 30. After a 5 cm push mid-task, re-solving the remaining program completes 23 of 60 pushed episodes and replaying the demonstration 6.

The same program executed on Panda, KUKA IIWA, Kinova Gen3 and UR5e.
One program executed on four arms without retraining (episode 227, side view). The episodes take 100, 99, 103 and 162 steps.
Three sampling seeds of one program, each with a different grasp.
Execution is a free variable. One effect, three grasps: seed 2 takes the bowl from the other side, and all three succeed.
Distance of the bowl to its programmed end over time, with and without a push.
A 5 cm push at step 28. The bowl's distance to its programmed end. Re-solving the remaining program absorbs the push; replaying the demonstration's actions ends 8.0 cm short.

A real arm programmed by a person

A person performs each task once at a table in front of a Franka Research 3 while a calibrated RGB-D camera records. One click per object on the first frame prompts SAM 2, whose masks and depth give each object's keypoints; the arm then executes the program closed loop from the same camera. The model is the simulation checkpoint fine-tuned on 156 real demonstrations of other tasks, none of the four evaluated. Over four tasks PEWAM completes 36 of 40 trials, against 22 for the same backbone conditioned on the video's last frame as a goal image. With the terminal state shifted 10 cm to either side the arm completes 9 and 9 of 10, and after a 5 cm push 7 of 10.

Task (successes of 10)Goal imagePEWAM
Cube to plate810
Cube into box510
Cube onto tray58
Box into bin48
Total (of 40)2236
A real Franka arm programmed by human videos. Each program is compiled from one human video; no robot demonstration of these tasks is used.
Programs drawn on the first frame of four human videos beside each video's last frame.
Programs compiled from real human videos. Each program is drawn on the demonstration's first frame, before the hand enters, beside the demonstration's last frame. Programs 1 and 2 share a scene: the same cube carried to a plate and to a box.
Ten further programs compiled from human videos.
Programs compiled from ten further real human demonstrations. Each panel draws the program on the first frame of the demonstration; the inset shows the last frame of the same demonstration.

What makes a program programmable

A model trained on the nine affine weights that define the contact points, the same information, reads every commanded rotation with a negative sign: contact has to be written as points, not weights. Without the paired corpus of several executions per effect, the model follows contact edits less closely.

ChangedMeasurePEWAMVariant
No contact channelfour arms, bowl / bottle89 / 4364 / 10
No paired corpuscontact following0.680.31
Demo-clock pointerheld-out bottle / Σ12 / 401 / 27
Native proprioceptionfour arms, bowl8928
Relative effectheld-out Σ401
Core ablations: one component changed per row, on the measure it acts on.

Citation

@article{hu2026pewam,
  title   = {Keep the Effect, Drop the Actor: Programmable Effect-to-Execution World-Action Models},
  author  = {Hu, Junyi and He, Zhewen and Fang, Yi},
  journal = {arXiv preprint},
  year    = {2026}
}