A post-training recipe that turns the public nvidia/Cosmos3-Edge world model into a robot-arm policy for the LIBERO-10 benchmark: ten long-horizon kitchen and tabletop tasks with a Franka arm in simulation. The recipe is proposed upstream in NVIDIA/cosmos-framework PR #278. The trained model is public: ubr-physical-ai/cosmos3-edge-libero10.
| # | Instruction | Success |
|---|
Two tasks pull the average down: putting both the alphabet soup and the tomato sauce in the basket (17/50), and placing both moka pots on the stove (25/50). Both need two pick-and-place moves of similar objects in one episode.
The same evaluation run on the public nvidia/Cosmos3-Edge checkpoint with no LIBERO training: same policy server, sampler, cameras and simulator settings. NVIDIA has not published a LIBERO number for Edge; this is our measurement. The released checkpoint has a trained action head for some robots, but its LIBERO slot is still at its initial values (never trained by NVIDIA), so this measures an untrained LIBERO output: a floor, not zero-shot skill.
Published LIBERO-10 (LIBERO-Long) success rates of other fine-tuned policies. Every row except the vanilla Cosmos3-Edge is a model post-trained on LIBERO by its authors; the vanilla row has an untrained LIBERO action head. Inputs, training data and trial counts differ, so this is context, not a controlled comparison.
| Model | Size | LIBERO-10 | Source |
|---|
Loss logged every iteration (faint) and its 25-iteration moving average (solid). The learning rate warms up over the first 500 iterations.
The upstream recipe runs on two 8-GPU nodes (HSDP 2 × 8, grad-accum 1). We kept the same experiment, learning-rate schedule and global batch and changed only the parallelism to fit four GPUs.
action_policy_server_libero on the final checkpoint, quantile-rot6d action normalization with the bundled LIBERO stats, 30 sampling steps, 20 fps.closed_loop_eval.py, agent-view + wrist cameras at 256 px, 50 trials per task, 8 parallel environments, seed 0.python cosmos_framework/simulation/libero/closed_loop_eval.py \ --server_url http://localhost:8000 --task_suite libero_10 \ --num_trials_per_task 50 --num_envs 8 \ --camera agentview,wrist --image_size 256 \ --action_space frame_wise_relative --rotation_space 6d --action_dim 10
Setup note: the LIBERO simulator environment pulls in egl-probe, which does not build with CMake 4; setting CMAKE_POLICY_VERSION_MINIMUM=3.5 fixes it.