MotorMind.
MotorMind

Scaffolding General Vision Language Models
for Zero-Shot Robot Manipulation

University of Illinois Urbana-Champaign

Equipping a VLM to operate robots like a human teleoperator

Reasoning directly from observations, issuing actions, and continuously adapting to execution feedback.

  • No task-specific training
  • No VLA-as-Tool
  • No Action Experts
  • No Coding Agent
  • No Motion Planning Tool or IK Module
  • No SAM3
  • No Skills/API
Explore the research

MotorMind bridges VLM reasoning and robot control through mid-level actions, continual refinement, and asynchronous scheduling, enabling strong zero-shot manipulation in both simulation (Franka in LIBERO-PRO) and on a real robot (xArm6), without task-specific training.

MotorMind connects visual observations, VLM reasoning, robot actions and feedback.
A frozen VLM proposes mid-level actions; the harness connects decisions to execution and measured feedback. Robot frames show an actual rollout; harness outputs and the interruption timeline are illustrative.

Understanding VLM decisions

Decision Diagnostic

Can a general-purpose VLM make the decisions needed for manipulation? This diagnostic evaluation probes three capabilities with 240 visual questions, 80 per capability, drawn from robot observations.

01

Action selection

What should the robot do next to advance the stated subgoal?

02

Progress assessment

Did the observed transition make progress toward the subgoal?

03

Subgoal completion

Does the current state satisfy the subgoal and its success criterion?

Diagnostic Results

Accuracy by decision type

Percent correct over 80 questions each. Action selection is the hardest task for every model.

Accuracy against latency

Overall accuracy on all 240 questions versus mean inference time per query (log scale).

Action Selection: Detailed Results
Action-selection accuracy by primitive type. Each cell reports accuracy (%) and correct answers / total questions.
ModelTranslation (48)Rotation (19)Gripper (13)
Qwen3.8-Flash-Next-FP847.92% (23/48)0.00% (0/19)46.15% (6/13)
HY-Embodied-0.5 MoT-2B20.83% (10/48)15.79% (3/19)23.08% (3/13)
Hy-Embodied-VLM-1.0 A3B31.25% (15/48)0.00% (0/19)0.00% (0/13)
Cosmos3-Nano (understanding tower)27.08% (13/48)0.00% (0/19)15.38% (2/13)
GLM-5.3-Flash41.67% (20/48)15.79% (3/19)53.85% (7/13)
GPT-6 Astra (medium)58.33% (28/48)63.16% (12/19)61.54% (8/13)

These decisions motivate the MotorMind design: propose short action sequences, observe their effects, and reassess progress using robot feedback.

From reasoning to action

How MotorMind Works

One frozen VLM takes five complementary roles: Planner, Executor, Monitor, Verifier, and Memory. A deterministic Controller turns its proposals into physical motion.

MotorMind architecture: planning, short action batches, control, background monitoring, verification, and memory.
01

Mid-level actions

The VLM proposes parameterized moves, rotations, and gripper actions. The controller validates and executes them.

02

Continual refinement

Fresh observations and measured feedback inform the next proposal. Outcomes determine whether to advance, retry, or replan.

03

Concurrent feedback

The Monitor checks ongoing execution and can cancel pending commands at an action boundary. Memory updates run in the background.

Zero-shot manipulation

Main Experiments

On LIBERO-PRO, MotorMind achieves the highest success among the evaluated zero-shot methods. Success rate and episode wall time are reported together.

Base task success: CaP-X 13.3%, MotorMind 66.7%, fine-tuned OpenVLA-OFT 98.3%, and fine-tuned pi 0.5 98.3%.
Perturbed task success: CaP-X 19.2%, MotorMind 53.8%, fine-tuned OpenVLA-OFT 51.2%, and fine-tuned pi 0.5 63.3%.

Success against episode wall time

Every configuration in the full results table, for the selected split. The dashed curve is equal cost-efficiency to MotorMind.

Full Results · View complete table· Hide complete table
LIBERO-PRO: success rates (%), episode wall time, and time-normalized success. Methods are grouped by whether the underlying policy uses task-specific fine-tuning.
MethodConfigurationBase success rate (%)Perturbation success rate (%)Base efficiencyPerturbation efficiency
GoalSpatialObjectAvg.SemanticObjectPositionTaskAvg.Time (s) ↓Time score ↑Time (s) ↓Time score ↑
Fine-Tuned Policies / Agentic Methods with Fine-Tuned Policies
π₀.₅Fine-tuned95.0100.0100.098.398.395.036.723.363.35.9997.16.8562.1
MolmoAct2Fine-tuned100.0100.0100.0100.095.090.036.730.062.96.6914.67.7492.2
OpenVLA / OFTFine-tuned95.0100.0100.098.396.786.711.710.051.25.71033.56.7459.9
GR00T N1.5Fine-tuned0.095.00.031.730.030.01.716.719.615.6121.816.272.4
VoLoAgent (FT VLA as Tool)Zero-shot agent55.090.0100.081.788.383.345.025.060.4106.246.1216.916.7
Harness VLA (FT VLA as Tool)Zero-shot35.030.030.031.730.031.728.325.028.8354.65.4342.55.0
Harness VLA (FT VLA as Tool)Learnt on task · 5 loops55.050.055.053.365.051.750.050.054.21040.43.11286.22.5
Harness VLA (FT VLA as Tool)Learnt on task · 10 loops70.090.080.080.075.066.753.366.765.41764.02.72184.11.8
Harness VLA (FT VLA as Tool)Learnt on base · 5 loops————55.048.343.341.747.1——1125.22.5
Harness VLA (FT VLA as Tool)Learnt on base · 10 loops————70.058.341.745.053.8——1866.21.7
Zero-Shot Methods
π₀.₅Zero-shot0.00.00.00.00.00.00.00.00.08.30.08.20.0
MolmoAct2Zero-shot0.00.00.00.00.00.00.00.00.09.40.09.30.0
OpenVLA / OFTZero-shot0.00.00.00.00.00.00.03.30.834.70.034.41.5
GR00T N1.5Zero-shot0.00.00.00.00.00.00.00.00.06.30.06.20.0
CaP-X (GPT5.6-Terra + Claude Opus 5)Zero-shot5.00.05.03.315.03.31.73.35.868.52.972.44.8
CaP-X (GPT5.6-Terra + Claude Opus 5)5 loops25.00.015.013.325.016.716.715.018.3290.12.8270.24.1
CaP-X (GPT5.6-Terra + Claude Opus 5)10 loops25.00.015.013.326.716.716.716.719.2346.42.3320.43.6
VoLoAgent (zero-shot VLA as Tool)Zero-shot0.00.00.00.00.00.00.00.00.0500.00.0500.00.0
MotorMind (Qwen3.8-Flash-Next)Zero-shot45.075.080.066.758.346.751.758.353.8223.417.91248.512.99

Time score = 60 × success rate (%) / episode wall time (s), in percentage points per minute. Bold and underlined values mark the best and second-best results within each group. A dash denotes an unreported result.

Backbone comparison

Swapping in a stronger backbone

One seed per suite, LIBERO-PRO base. GPT-6 Sol with medium reasoning, same harness.

Average success rises from 66.7% to 83.3%, with the largest gains on Object and Goal (+20 points each). Every suite takes longer; Spatial goes from 199.0 s to 370.4 s.

When the world changes

Adapting During Execution

Moving objects, shifted destinations, and revised instructions require decisions grounded in the latest observation.

Dynamic Reasoning

70% · 10 tasks

Identify moving targets through semantic, relational, or temporal descriptions.

Scene Shift

90% · 10 tasks

Continue toward the same goal after an object or destination moves.

Dynamic Manipulation

80% · 5 tasks

Track and manipulate targets moving on a conveyor.

Prompt Shift

60% · 5 tasks

Respond to an updated instruction without resetting the scene.

All 30 adaptive task instructions

From simulation to the real world

Real Robot Deployment

The same VLM-facing interface operates an xArm6 with RealSense D455 observations, without task-specific demonstrations or policy fine-tuning.

95%

Success across direct manipulation
and human perturbations · 114 / 120 trials

Observe. Reassess. Recover.

When an approach heads toward the wrong region, the Monitor flags the mismatch. The Executor uses a new observation to re-localize the target and continue.

Case study: recovering from a wrong-object approach

xArm6 corrects its initial approach and places a blue cube into the box.
Execution and recovery on the real robot. The pooled 95% result covers direct manipulation and human perturbations; semantic-understanding trials are reported separately in the paper.

Results by object and instruction

Semantic Understanding

Success over 5 trials per task.

Failure analysis

Where the remaining errors come from

We labeled every failure in the 300 LIBERO-PRO episodes behind the MotorMind rows of the main table, using the simulator ground truth the evaluation logged at every observation: each object's pose, the finger gap and each goal predicate. None of it reaches a model. Failures are labeled along two axes.

Physical failures describe what happened in the scene: an empty grasp, a wrong object lifted, the target dropped, knocked over, or misplaced off its goal. A failure is recovered if the target is carried again afterwards; a failed episode with no unrecovered physical failure is stuck.

VLM decision errors are model judgments that the simulator state contradicts: in grounding, planning, progress checks, task completion, or failure monitoring. Only model-made verdicts are scored; verdicts the harness settles in code are excluded.

From failure to outcome

Perturbation lowers the share of failure-free episodes from 57% to 44%. Recovery depends on the failure type, and false done occurs only under perturbation (21 of 240 episodes).

Each failed episode is assigned its last unrecovered failure. Band thickness is the share of episodes in the split.

Recovery by physical failure type

Failure events across all 300 episodes (an episode can have several), and the share after which the target was carried again.

Grasp acquisition is the main bottleneck

63 of the 131 failed episodes (48%) never get hold of the target: 31 are stuck without lifting it, 27 of them without ever closing the jaws, and 32 end after empty grasps.

In 42 of those 63 the target is a bowl or a plate, which offers only a thin rim to grasp or has to be pushed. Tasks with such targets succeed 47% of the time (67/144), against 70% (95/136) for other objects.

Of the 54 stuck episodes, 31 never lift the target, 16 run out of budget while holding it or after carrying it, and 7 never achieve an articulation goal such as opening a drawer.

Decision errors: grounding and premature “done”

VLM decision errors per 60 episodes

Raw counts by perturbation. Malformed replies are counted separately.

Grounding errors lead to wrong-object lifts

A grounding error is counted when the model localizes a target point near an unrelated object and far from the true target during an approach or grasp. 23 of the 25 wrong-object lifts follow a grounding error in the same or an earlier subgoal. Episodes with a grounding error succeed 31% of the time (14/45), against 61% (155/255) without one.

Step checks are rarely wrong; task checks often are

Across 300 episodes there are only 12 step-level errors: 4 progress checks, 5 missed failures and 3 planning errors. The Planner phrases step criteria in measured quantities, such as whether the gripper reports holding and how far the tool rose, so the Verifier checks readings instead of inferring progress from pixels. Task completion has no such reading and remains a visual judgment.

All 71 “task done” claims were wrong

Episodes stop as soon as the goal predicate holds, so every claim was made while it was false. We checked all 71 against the recorded state.

What the scene looked like at the claim

What the harness did with it

Citation

@article{li2026motormind,
  title = {MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation},
  author = {Li, Bingxuan and Song, Siqi and Wu, Yizhuo and Yao, Jiarui and Zhang, Tong and Zhang, Huan},
  journal = {arXiv preprint arXiv:2609.38078},
  year = {2026},
  eprint = {2609.38078},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.38078}
}

Figure explorer

100%Original ↗

Zoom to inspect details · Drag to pan · Use ← / → to browse · Esc to close