A language goal becomes a visual choice, an inspected pose, and a robot action.
RobotUse connects language-model decisions to physical execution through images and geometry. The agent points to a target, chooses a gripper pose, checks the fresh view, and decides whether to continue or revise. The backend realizes those choices and returns the next observation.
Language–visual handoff
The handoff carries the physical subgoal and its visual evidence. Each local choice stays connected to the scene that prompted it and the robot motion it produces.
Native camera projections and point-cloud views connect each decision to the scene. The simulator replay retains the arm motion; physical replays pair the same interface with recorded camera observations.
How the main agent and subagent share context
The main agent retains the task goal and returned reports. The subagent retains its observations, selections, pose decisions, and execution feedback. A report returns the chosen action or current outcome with the visual evidence needed for the next task-level decision. Local roles are grouped under “Subagent” in this presentation.
Learning from execution
A persistent playbook carries guidance from earlier attempts. This example shows two separate bowl-stacking runs: both pick up the bowl, while the later run carries it into a supported placement.
Bowl placement across two attempts
Selected qualitative cases. Candidate choices and execution conditions also differ; the videos are separate complete runs.
Physical robot demonstrations
The visual action interface was adapted to a physical Franka Panda. These two examples show the robot moving a cube into a tray and stacking it on another cube.
Pick & place
Stack cubes
Camera observation replays from physical executions. Final frames show the requested object relations. These demonstrations are selected execution records and do not establish an aggregate physical-robot success rate.
Results on RoboLab
Evaluation spans 40 manipulation tasks: 21 Simple, 13 Moderate, and 6 Complex. RobotUse completes 54 of 120 episodes, achieving 45.00% task success.
| Method | Overall | Simple | Moderate | Complex | Episodes |
|---|---|---|---|---|---|
| RobotUse | 45.00 | 53.97 | 35.90 | 33.33 | 120 |
| CaP-X | 38.33 | 42.86 | 35.90 | 27.78 | 120 |
| Open Robot Skill | 20.00 | 31.75 | 10.26 | 0.00 | 120 |
| mini-SWE | 20.83 | 25.40 | 17.95 | 11.11 | 120 |
| π₀.₅ | 30.75 | 30.48 | 33.08 | 26.67 | 400 |
| Cosmos 3 | 42.25 | 44.76 | 45.38 | 26.67 | 400 |
Language agents use three trials per task; direct-action policies use ten. Full-task success uses the final task verifier. Execution limits and selected rerun protocols differ by method. Qualitative videos above are selected separately from these aggregate results.
Evaluation sources and conditions
RobotUse, CaP-X, and Open Robot Skill aggregates come from the completed 40-task benchmark reports. Revision results use 40 selected episodes per revision. The π₀.₅ and mini-SWE results were recomputed from episode records; Cosmos 3 uses the supplied 40-task report. Different stopping rules and episode counts make this a system comparison.
Code & getting started
The public release includes the core agent harness, playbooks, visual action tools, and native RoboLab integration. The optional web UI is a separate entry point. The physical Panda adapter is not included in this release.
git clone --recurse-submodules https://github.com/robotuse-team/RobotUse.git
cd RobotUse
scripts/setup/sources.shStart with the setup guide for Git LFS, the three core environments, model credentials, and simulator assets. Linux and an NVIDIA RTX GPU are required for simulation.
RobotUse code is Apache-2.0. Some upstream model and simulator assets carry separate noncommercial restrictions. Media provenance records episode identifiers, playback representation, and file hashes.