第 6 章 duck-sim:没有机器人也能开发

本章代码走读:scripts/duck-sim(shell 脚本头注释)、docs/design/simulation.md。

6.1 一条命令一只鸭

scripts/duck-sim 是个 shell 脚本,它的头注释就是完整的用户手册:

#!/bin/sh
# A duck in MuJoCo, driven by the real daemon, in one command.
#
#   scripts/duck-sim              # a window opens, the duck stands up, and it is yours
#   scripts/duck-sim drive        # walk it forward for a few seconds
#   scripts/duck-sim ctl health   # anything robotctl does
#   scripts/duck-sim down
#
# And the same duck, as a machine you log into:
#
#   scripts/duck-sim boot         # the daemons under real systemd, in a container
#   scripts/duck-sim shell        # you are on the duck
#   duck-a # robotctl health
#
# `boot` needs sudo, because `systemd-nspawn` does. What it buys over the plain form is the part
# that is not fakeable: the real unit files, with their real User=, groups, RuntimeDirectory= and
# hardening, under a real init — which is what makes `robotctl update apply` and the health gate
# behave here the way they behave on a robot.

两级真实度,按需升级:普通模式起 MuJoCo 窗口与真实 robotd;boot 模式把守护进程装进 systemd-nspawn 容器,跑真实的 systemd 单元文件——真实的 User=、用户组、RuntimeDirectory= 与加固选项。注释说得很直白:单元文件的权限与沙箱配置"不可伪造",而正是它们决定 robotctl update apply 和健康门禁在仿真里是否与真机同构。容器还给了鸭子一个主机名(duck-a),四只鸭子各自登录、互相发现——多机系统的调试体验也有了。

6.2 设计文档:一只"孪生鸭"的边界

docs/design/simulation.md 开篇给出定位与实测数据:

**Status:** built and in use. `robotd --sim`, `tofd --sim` and `mediad --sim-camera` are here, with
`microduck_rl`'s `duck-body` serving the other half; the containers (§8), the ether (§5) and the
per-duck cameras all run from `scripts/duck-sim`. Measured: the daemon holds `50.0 of 50.0 Hz · 0
missed` against a MuJoCo body, detects a seated boot from the simulator's own joint angles, and
`robotctl robot init` runs the sitstand policy until the duck is upright and stays there; four ducks
in containers sing a full chorale over the ether. ...

The goal is a duck you develop against exactly as you develop against a robot: the same binaries,
the same units, the same `robotctl`, the same `duckctl open` — with the body in MuJoCo instead of on
the desk. Not a mock, and not a test harness. A twin, with a written-down boundary.

关键事实:

6.3 代码走读:三个"没人猜得出来"的坑

脚本头注释的最后一段,把脚本存在的理由总结为堵住三个坑(每一条都是环境差异的化石):

# The three things this exists to stop anyone typing by hand, because none of them is guessable:
#
#   * policies live at /opt/robot/daemon/current on a robot and in this repo on a laptop, so every
#     one has to be named in a params file;
#   * `ort` dlopens libonnxruntime and a laptop has no system one — the RL repo's venv does;
#   * a unix socket path is capped at ~108 bytes, so the state directory has to be short.
  1. 策略路径双轨:机器人上在 /opt/robot/daemon/current,笔记本上在仓库内——于是仿真参数文件必须逐个点名策略;
  2. ONNX Runtime 的动态加载(第 1 章伏笔回收):robotd dlopen 系统 libonnxruntime,笔记本通常没有——但 RL 仓库的 venv 里有,脚本去那里找。开发环境复用训练环境,一个实际的"仓库间依赖";
  3. Unix socket 路径约 108 字节上限:每服务一 socket(第 3 章)的架构下,状态目录必须短——这是内核限制倒逼的目录命名设计。

把"不可猜的坑"封装进一条命令,是开发者体验(DX)工程的核心:不是文档写得细,是让人根本不需要读文档。

6.4 与 RL 训练仿真的分工

duck-sim(主仓库) mjlab 环境(RL 仓库)
目的 软件在环:验证守护进程、协议、部署链路 大规模并行训练(4096 环境)
运行的代码 真实 Rust 守护进程(--sim 开关) 纯 Python 训练环境
身体服务方 RL 仓库的 duck-body mjlab/MuJoCo Warp
频率 实时,实测 50.0/50.0 Hz 零丢拍 GPU 加速,非实时
逼真度侧重 系统行为(systemd、更新、门禁) 执行器/齿隙等物理保真

两个仿真不是竞争关系而是一条验证链的两段:RL 仓库训出策略 → duck-sim 验证"真守护进程吃得下" → 真机。RL 仓库的 scripts/ 里还有专门的 sim2real 对比脚本,把同一条策略在两侧的轨迹并排量差。

6.5 无硬件学习路线(修订版)

结合 Sandbox 与 duck-sim,零硬件的完整路径:

  1. 第 0 天:Hugging Face Sandbox 浏览器里建立手感;
  2. 第 1 周:scripts/duck-sim 起鸭,robotctl 遍历命令;再 boot 进容器看 systemd 单元与健康门禁的真实行为;
  3. 第 2–3 周:读 duck-ipc-proto 与 robotd/src/control.rs,在仿真里给自己的 RPC 方法或技能写代码——改 control.rs 是安全的,它无 IO、可单测;
  4. 第 4 周起:进入 RL 仓库(下一章),训练并用 duck-sim 验证自己的策略。

6.6 本章小结

duck-sim 的启示不在 MuJoCo,而在三件纪律:孪生要写明边界、真实度分级可选(普通/容器)、把环境差异的坑封装成一条命令。"同一套真软件,两种身体"是第 3 章契约设计的最大红利,也是这个仓库对"机器人开发难"给出的最系统性的回答。至此工程篇结束,下半部进入策略的产地。