strands-isaaclab-go2-rough-policy

rsl_rl PPO policy for Unitree Go2 quadruped rough-terrain locomotion (Isaac-Velocity-Rough-UnitreeGo2), trained with strands-robots' isaaclab train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-go2-rough.

playback: 4 parallel envs

envs 0–3 (2×2) from the strands camera — mp4.

Results

parallel envs × iterations 4096 × 1500 (147 M env steps), PhysX (physics=isaacsim_physx)
wall time 63 min 38 s on 1× NVIDIA L40S
env-steps/s median 43 k, max 66 k (GPU shared with other runs most of the time)
mean reward -0.76 → 31.4 best → 30.6 last
velocity-tracking success 0.98 last iteration; mean episode length 995 of 1000; terrain level 5.8
recorded rollout 12 / 12 envs walked the full 10 s window; mean return 19.6
PPO iteration mean reward velocity-tracking success mean ep. length (of 1000) terrain curriculum level env-steps/s
0 -0.76 0.042 12 3.51 33,728
150 7.78 0.176 889 0.32 50,092
375 16.37 0.610 960 1.41 40,384
750 24.35 0.923 990 5.74 43,301
1125 28.10 0.990 981 5.93 42,916
1499 30.59 0.979 995 5.80 43,387

Full per-iteration metrics: train_curve.json.

Files

file what
model_1499.pt final rsl_rl checkpoint (OnPolicyRunner.load) — actor + critic + optimizer
exported/policy.pt TorchScript actor incl. observation normalizer (deterministic mean action); input (N, 235) → (N, 12)
exported/policy.onnx (+ .onnx.data) the same actor as ONNX
params/env.yaml, params/agent.yaml exact Isaac Lab env + rsl_rl agent configs of the run (seed 1)
train_curve.json · record.json · verify.json training curve · recording metadata (obs/action names) · dataset verification
examples/record_trained_policy.py the strands recording script used for the dataset
playback.mp4/.gif, frame.png media

How it was made with strands-robots

Setup

# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
  --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES          # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer"   # strands-robots with PR #4227

Train

As an agent tool call (the train_policy tool is a Strands @tool):

from strands import Agent
from strands_robots.tools.train_policy import train_policy

agent = Agent(tools=[train_policy])
agent("Train the Unitree Go2 to walk on rough terrain with the isaaclab provider: task Isaac-Velocity-Rough-UnitreeGo2, 4096 envs, PhysX, 1500 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c4_go2_rough",
#                 extra={"task": "Isaac-Velocity-Rough-UnitreeGo2", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})

As plain Python (exactly what produced this run):

from strands_robots.tools.train_policy import train_policy

job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1,
                   output_dir="runs/c4_go2_rough",
                   extra={"task": "Isaac-Velocity-Rough-UnitreeGo2", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})
# poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")

Under the hood the provider runs python -m isaaclab train --rl_library rsl_rl --task Isaac-Velocity-Rough-UnitreeGo2 --max_iterations 1500 --num_envs 4096 --seed 1 physics=isaacsim_physx in $ISAACLAB_PYTHON and parses its log. Job id of this run: isaaclab-20260929-053304-229d959f49c7. Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.

Record (dataset repo)

The final checkpoint was rolled out and recorded with examples/record_trained_policy.py (included in this repo). It runs in the Isaac Lab venv with strands on PYTHONPATH and:

  1. rebuilds the task env in play mode and adds an RTX camera per env;
  2. loads model_1499.pt with rsl_rl's OnPolicyRunner and exports TorchScript/ONNX with Isaac Lab's exporter;
  3. wraps the exported actor as a strands Policy (RslRlJitPolicy, max |Δa| vs rsl_rl inference = 2.4e-07);
  4. steps the env with policy.get_actions_sync(...) and writes every frame through strands DatasetRecorder (strands_robots.dataset_recorder.DatasetRecorder.create(...) → add_frame → save_episode, LeRobot v3);
  5. verifies the result with strands verify_dataset + LeRobotDataset load + video decode + NaN scan.
OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
  --task Isaac-Velocity-Rough-UnitreeGo2 --checkpoint model_1499.pt --episodes 12 --frames 500 \
  --cam chase --override physics=isaacsim_physx \
  --task_str "walk over rough terrain following the commanded base velocity" --robot_type unitree_go2 \
  --root out/ds --repo_id cagataydev/strands-isaaclab-go2-rough --attach Robot/base --eye 1.0,-1.5,0.45
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-go2-rough

Use it

Play it in Isaac Lab (in the Isaac Lab venv):

huggingface-cli download cagataydev/strands-isaaclab-go2-rough-policy --local-dir go2_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Velocity-Rough-UnitreeGo2 \
  --num_envs 16 --checkpoint go2_policy/model_1499.pt physics=isaacsim_physx   # PhysX! (IL-X-011); add --video --video_length 500

Raw TorchScript actor:

import torch
pi = torch.jit.load("go2_policy/exported/policy.pt").eval()
actions = pi(obs)          # obs: (N, 235) concatenated Isaac Lab 'policy' observation group -> (N, 12)

As a strands Policy. create_policy("rl") cannot load rsl_rl checkpoints yet (IL-X-006: strands' RL actor is a Tanh MLP, rsl_rl's is ELU with a baked-in normalizer), so wrap the exported actor — this is the adapter the recording script uses:

import numpy as np, torch
from strands_robots.policies.base import Policy

class RslRlJitPolicy(Policy):
    def __init__(self, jit_path, action_names, device="cpu"):
        self.net = torch.jit.load(jit_path, map_location=device).eval()
        self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
    @property
    def provider_name(self): return "isaaclab_rsl_rl_jit"
    def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
    def reset(self, seed=None): self.net.reset()
    async def get_actions(self, observation_dict, instruction, **kw):
        x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
        with torch.inference_mode():
            y = self.net(x.unsqueeze(0))[0].cpu().numpy()
        return [dict(zip(self.action_names, map(float, y)))]

import json
names = json.load(open("go2_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("go2_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_235}, "walk forward")   # [{"joint_pos.FL_hip_joint": ..., ...}]

The observation has to come from the Isaac Lab task (base velocities, gravity, velocity command, joint states, last action, height scan …), so the policy runs inside Isaac Lab; see examples/record_trained_policy.py for the full env + camera + DatasetRecorder loop.

Provenance

  • strands-robots: feat/isaaclab-trainer @ fa66fc68 — strands-labs/robots#4227 (isaaclab train_policy provider, IsaacLabTrainer; DatasetRecorder; verify_dataset)
  • Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · PhysX (physics=isaacsim_physx) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA
  • GPU: 1× NVIDIA L40S (46 GB), shared with the H1 rough-terrain run for all of training
  • Seeds: training seed 1 (params/agent.yaml, params/env.yaml); recording seed 7
  • Training job: isaaclab-20260929-053304-229d959f49c7, 2026-09-29

Limitations

  • Simulation only. Nothing here was run on a real Unitree Go2; no sim-to-real claims (the policy was not trained with sim-to-real hardening beyond Isaac Lab's default randomization).
  • Release candidates: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ.
  • Physics preset matters (IL-X-011): trained and recorded on PhysX only; always pass physics=isaacsim_physx when playing it (the provider does not remember the preset; the G1 PhysX policy falls in < 1.1 s when replayed on Newton).
  • Recording: 12 parallel envs from a common reset, 500 frames (10 s) each with a chase camera; all 12 ran the full window. observation.state mixes joint positions with the policy observation (incl. height scan), so it is wide (254-D) and not a standard LeRobot "robot state".
  • create_policy("rl") in strands cannot load this rsl_rl checkpoint yet (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU
    • obs-normalizer); use the exported TorchScript + the small wrapper shown above.
  • Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback).

License

Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:

  • The recorded trajectories, rendered camera video, playback clips and the trained policy weights are user-generated content produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as isaacsim/LICENSE.txt) §2.1 explicitly allows you to "distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release that content under CC-BY-4.0.
  • No NVIDIA Content is redistributed: the Unitree Go2 USD (IsaacLab/Robots/Unitree/Go2/go2.usd) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are not in this repo; params/env.yaml only references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. The rough terrain is procedurally generated by Isaac Lab. "Unitree Go2" is a product of Unitree Robotics; no endorsement by Unitree or NVIDIA is implied.
  • params/*.yaml are Isaac Lab task / agent configurations (Isaac Lab is BSD-3-Clause); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train cagataydev/strands-isaaclab-go2-rough-policy

Collection including cagataydev/strands-isaaclab-go2-rough-policy

Evaluation results

  • velocity-tracking success rate (last training iteration) on Isaac Lab Isaac-Velocity-Rough-UnitreeGo2 (4096 envs, PhysX)
    self-reported
    0.979
  • mean episode reward (last iteration) on Isaac Lab Isaac-Velocity-Rough-UnitreeGo2 (4096 envs, PhysX)
    self-reported
    30.590