RobotUseResearch project · 2026

RobotUse: Allocating
Computation, Context, and Decisions

A language–visual interface for choosing and inspecting robot actions.

1 KAIST · 2 Seoul National University* Equal contributionarXiv ↗Code ↗Get started ↓

How RobotUse works

Follow an instruction into a visual choice, then watch the Subagent’s report return to the Main agent.
Language → visual choice → robot action

Follow the main agent’s instruction into a point, a gripper pose and robot action.

Stack the left bowl on the right bowl.

01

Main agent

Sent instruction → Subagent
Get a clear overhead view of both bowls.
→
02

Subagent

Visual choice in progress
Choose a view that shows both bowls.
→
03

Backend

Visual feedback & execution
Return the scene and realize the requested viewing pose.
Main agent → Subagent · physical subgoal
Recorded simulator motionFranka Panda · parallel-jaw gripper

Native target selectionPoint + mask
Native visual action evidence
Open original tool image ↗
Recorded instruction & report

Trace & timing · Download this demo ↓.

A language goal becomes a visual choice, an inspected pose, and a robot action.

RobotUse connects language-model decisions to physical execution through images and geometry. The agent points to a target, chooses a gripper pose, checks the fresh view, and decides whether to continue or revise. The backend realizes those choices and returns the next observation.

1.

Language–visual handoff

The handoff carries the physical subgoal and its visual evidence. Each local choice stays connected to the scene that prompted it and the robot motion it produces.

Native camera projections and point-cloud views connect each decision to the scene. The simulator replay retains the arm motion; physical replays pair the same interface with recorded camera observations.

How the main agent and subagent share context

The main agent retains the task goal and returned reports. The subagent retains its observations, selections, pose decisions, and execution feedback. A report returns the chosen action or current outcome with the visual evidence needed for the next task-level decision. Local roles are grouped under “Subagent” in this presentation.

Learning from execution

A persistent playbook carries guidance from earlier attempts. This example shows two separate bowl-stacking runs: both pick up the bowl, while the later run carries it into a supported placement.

Bowl placement across two attempts

Earlier executionNative failure
Later executionNative success

Selected qualitative cases. Candidate choices and execution conditions also differ; the videos are separate complete runs.

2.

Physical robot demonstrations

The visual action interface was adapted to a physical Franka Panda. These two examples show the robot moving a cube into a tray and stacking it on another cube.

Pick & place

Orange cube into the tray.12 camera observations

Stack cubes

Orange cube onto the green cube.11 camera observations

Camera observation replays from physical executions. Final frames show the requested object relations. These demonstrations are selected execution records and do not establish an aggregate physical-robot success rate.

3.

Results on RoboLab

Evaluation spans 40 manipulation tasks: 21 Simple, 13 Moderate, and 6 Complex. RobotUse completes 54 of 120 episodes, achieving 45.00% task success.

40Manipulation tasks
45.0%RobotUse success
+6.67Points over CaP-X
Full-task success on RoboLab (%)
MethodOverallSimpleModerateComplexEpisodes
RobotUse45.0053.9735.9033.33120
CaP-X38.3342.8635.9027.78120
Open Robot Skill20.0031.7510.260.00120
mini-SWE20.8325.4017.9511.11120
π₀.₅30.7530.4833.0826.67400
Cosmos 342.2544.7645.3826.67400

Language agents use three trials per task; direct-action policies use ten. Full-task success uses the final task verifier. Execution limits and selected rerun protocols differ by method. Qualitative videos above are selected separately from these aggregate results.

Evaluation sources and conditions

RobotUse, CaP-X, and Open Robot Skill aggregates come from the completed 40-task benchmark reports. Revision results use 40 selected episodes per revision. The π₀.₅ and mini-SWE results were recomputed from episode records; Cosmos 3 uses the supplied 40-task report. Different stopping rules and episode counts make this a system comparison.

Download the compact result table and source notes.

4.

Code & getting started

The public release includes the core agent harness, playbooks, visual action tools, and native RoboLab integration. The optional web UI is a separate entry point. The physical Panda adapter is not included in this release.

Clone & prepare sources
git clone --recurse-submodules https://github.com/robotuse-team/RobotUse.git
cd RobotUse
scripts/setup/sources.sh

Start with the setup guide for Git LFS, the three core environments, model credentials, and simulator assets. Linux and an NVIDIA RTX GPU are required for simulation.

RobotUse code is Apache-2.0. Some upstream model and simulator assets carry separate noncommercial restrictions. Media provenance records episode identifiers, playback representation, and file hashes.