UB Robotics · NVIDIA Cosmos Codefest 2026

Cosmos3-Edge on LIBERO-10

A post-training recipe that turns the public nvidia/Cosmos3-Edge world model into a robot-arm policy for the LIBERO-10 benchmark: ten long-horizon kitchen and tabletop tasks with a Franka arm in simulation. The recipe is proposed upstream in NVIDIA/cosmos-framework PR #278. The trained model is public: ubr-physical-ai/cosmos3-edge-libero10.

Closed-loop success, 500 episodes
81.4%
Successes
407 / 500
Checkpoint
iter 2000
Trials per task
50
Tasks above 85 %
7 of 10
Model
download (public)

The policy at work

One episode per LIBERO-10 task with the final checkpoint, agent-view camera, instruction on top. All ten of these succeeded; the success rate above comes from the separate 500-episode run.

Per-task results

#InstructionSuccess

Two tasks pull the average down: putting both the alphabet soup and the tomato sauce in the basket (17/50), and placing both moka pots on the stove (25/50). Both need two pick-and-place moves of similar objects in one episode.

Without fine-tuning

The same evaluation run on the public nvidia/Cosmos3-Edge checkpoint with no LIBERO training: same policy server, sampler, cameras and simulator settings. NVIDIA has not published a LIBERO number for Edge; this is our measurement. The released checkpoint has a trained action head for some robots, but its LIBERO slot is still at its initial values (never trained by NVIDIA), so this measures an untrained LIBERO output: a floor, not zero-shot skill.

Vanilla Cosmos3-Edge, 500 episodes
0.0%
Successes
0 / 500 (0 of 50 on every task)
Recorded run
0 / 10 (every episode hit the 520-step limit)
Fine-tuned
407 / 500 = 81.4 %
Same task and start state, one episode per task. Left: fine-tuned with this recipe. Right: the vanilla checkpoint, which never completes a task.

How it compares

Published LIBERO-10 (LIBERO-Long) success rates of other fine-tuned policies. Every row except the vanilla Cosmos3-Edge is a model post-trained on LIBERO by its authors; the vanilla row has an untrained LIBERO action head. Inputs, training data and trial counts differ, so this is context, not a controlled comparison.

ModelSizeLIBERO-10Source

Training loss

Loss logged every iteration (faint) and its 25-iteration moving average (solid). The learning rate warms up over the first 500 iterations.

The run

Base modelnvidia/Cosmos3-Edge (4B, Mixture-of-Transformers)
DataLIBERO-10, LeRobot v3
Hardware4 × RTX PRO 6000 Blackwell 96 GB
Wall time3 d 4 h, 2,000 iterations
Global batch2,048 (128 × 4 GPUs × grad-accum 4)
Optimizerlr 5e-5, 500 warm-up steps
Precisionbfloat16, FSDP shard 4
Actions10-D, frame-wise relative, rot6d
Weightsubr-physical-ai/cosmos3-edge-libero10 (iteration 2000, public)

The upstream recipe runs on two 8-GPU nodes (HSDP 2 × 8, grad-accum 1). We kept the same experiment, learning-rate schedule and global batch and changed only the parallelism to fit four GPUs.

How it was evaluated

python cosmos_framework/simulation/libero/closed_loop_eval.py \
  --server_url http://localhost:8000 --task_suite libero_10 \
  --num_trials_per_task 50 --num_envs 8 \
  --camera agentview,wrist --image_size 256 \
  --action_space frame_wise_relative --rotation_space 6d --action_dim 10

Setup note: the LIBERO simulator environment pulls in egl-probe, which does not build with CMake 4; setting CMAKE_POLICY_VERSION_MINIMUM=3.5 fixes it.