第 9 章 Sim-to-Real:执行器、齿隙与随机化的纪律

本章代码走读:actuator/friction_dr_bam.py(全文 112 行,两代执行器类)、robot/microduck/add_backlash.py、AGENTS.md 不变量清单。

RL 仓库开篇一句话立纲:"Sim2real transfer is the whole point"——本章的三段代码就是这句话的全部工程实现。

9.1 代码走读:FrictionDRBamActuator——给 BAM 开摩擦随机化的口子

"""BAM actuator with per-env friction-magnitude domain randomization.

The canonical ``bam.mjlab.BamActuator`` exposes per-env gain scaling (kp/kd) but
no friction hook, and under BAM MuJoCo's ``dof_frictionloss`` is zeroed in
``edit_spec`` (BAM computes friction itself in ``compute()``). So the stock
``dr.dof_frictionloss`` is a no-op here.

This thin subclass adds a per-env ``friction_scale`` that multiplies BAM's
velocity-INDEPENDENT friction budget (Coulomb + Stribeck + load-dependent) inside
``_compute_friction_budget`` — the term that carries the dominant sim2real
friction uncertainty (stiction / gearbox). The viscous (velocity-proportional)
term is left at nominal; scale it too by overriding ``compute`` if ever needed.

Non-accumulating: ``friction_scale`` is reset to 1.0 then set to a fresh sample
each episode by the ``randomize_bam_friction`` event (see tasks/mdp.py).
"""

三层信息:

  1. 静默 no-op 的陷阱:BAM(Black-box Actuator Model)在初始化时把 MuJoCo 的 dof_frictionloss 清零、改由自己算摩擦——所以拿现成的 dr.dof_frictionloss 域随机化去"随机化摩擦",什么都不发生,也不报错。这是域随机化最阴的故障形态:你以为在练鲁棒性,其实策略从没见过摩擦变化。AGENTS.md 把它列为不变量:"joint-friction DR must scale the actuator's friction_scale — dof_frictionloss is zeroed under BAM, so randomizing it is a silent no-op."
  2. 只随机主导项:摩擦预算里"速度无关"的部分(库仑 + Stribeck + 负载相关)承载了主要的 sim2real 不确定性(静摩擦/齿轮箱),按环境随机化;粘性项(速度比例)留名义值。随机化的刀口要落在不确定最大的量上。
  3. 非累积:friction_scale 每集先复位 1.0 再采新样——第 9.4 节的"累积事故"正是这条纪律的由来。

9.2 代码走读:BacklashEncoderBamActuator——编码器在齿隙的哪一侧

第二个类解决一个微妙得多的问题:舵机固件的 PD 环读到的角度,在齿隙的哪一侧?

class BacklashEncoderBamActuator(FrictionDRBamActuator):
    """FrictionDRBamActuator whose firmware PD reads the encoder THROUGH backlash.

    Backlash models (robot_groundcontact_backlash.xml) put an unactuated
    ``passive_<joint>_backlash`` hinge in series with each servo joint: the
    servo joint is the motor output, the backlash joint is the play between it
    and the link, and the link angle is their sum.

    On the real servo the magnetic encoder sits on the OUTPUT side of that
    play, so the firmware position loop closes on main+backlash — while the
    servo winds through the dead zone the measured position (and hence the PD
    error) doesn't change. This subclass reproduces that: ``cmd.pos`` fed to
    BAM's voltage control law becomes qpos[main] + qpos[backlash].

    ``cmd.vel`` is left motor-side on purpose: in BAM it drives back-EMF and
    friction, which are rotor physics, not an encoder-derived firmware signal.

    Degrades to a plain FrictionDRBamActuator on models without backlash
    joints (per-joint mask), so it is safe to use on any microduck model.
    """

代码本体只有一处改动,却精确复刻了真实舵机的控制结构:

    def get_command(self, data) -> ActuatorCmd:
        cmd = super().get_command(data)
        pos = cmd.pos + data.joint_pos[:, self._backlash_joint_ids] * self._backlash_mask
        return dataclasses.replace(cmd, pos=pos)

要点拆解:真实 XL330 的磁编码器在齿隙的输出侧,固件位置环闭合于"主关节 + 齿隙"之和——舵机在死区里转时,被测位置(进而 PD 误差)纹丝不动。仿真里把 cmd.pos 换成 qpos[main] + qpos[backlash],就再现了"电机转了、编码器没动"的死区行为。而 cmd.vel 故意留在电机侧:它驱动反电动势与摩擦计算,那是转子物理,不是编码器信号。同一行代码里两个量走不同侧——保真度建模要细到"每个信号在物理上来自哪里"。最后,对无齿隙模型按关节掩码优雅降级,一个类通吃全部模型。

9.3 代码走读:add_backlash.py——从 CAD 到 XML 的齿隙注入

齿隙在 MJCF 里的实现是每个舵机关节后串一个无驱动被动铰链:

"""Inject gearbox-backlash joints into an onshape-to-robot MJCF export.

For every actuated servo joint (``class="chosen_actuator"``) this inserts an
unactuated hinge on the same body / same axis right after it:

    <joint axis="0 0 1" name="left_hip_yaw" ... class="chosen_actuator"/>
    <joint axis="0 0 1" name="passive_left_hip_yaw_backlash" class="backlash"/>

The composite link rotation is main + backlash: the main joint is the servo
output (BAM drives it), the backlash joint is the play between the servo and
the link, free to wander within ±(backlash/2).

Naming: the ``passive_`` prefix means the new joints are automatically excluded
by every existing regex in the task configs (actuators ``^(?!passive_).*``,
joint obs, pose reward).

``--backlash-deg`` is the TOTAL peak-to-peak play (what you measure wiggling
the horn with the servo held); the joint range is symmetric ±deg/2.
"""

三处巧思:命名即排除——passive_ 前缀让所有现存正则(执行器、观测、姿态奖励的选择器)自动跳过新关节,注入脚本零改动兼容全仓库;参数以测量定义——--backlash-deg 2.0 是"握住舵机晃动输出喇叭测到的总峰峰值",关节范围取对称 ±1°,参数的物理含义直接对应一种手持测量操作;接触参数按仿真步长定制——默认 solref (0.02,1) 在这种小范围关节上会让限位超程约两倍,改用 0.01 = 2×sim_dt(dt=0.005 时"最 stiff 的稳定设置"),solimp 提高阻抗让齿面接触近乎刚性。

为什么对 ±1° 这么较真? 因为齿隙是"结构已知的误差",用它建被动铰链 + 编码器侧反馈,策略在训练时就体验过死区里"指令动了、反馈没动"的因果结构。对比第 8 章的 head droop:随机化对付的是"不知模型的残差",结构化建模对付的是"知道结构的误差"——能用物理结构表达的,不要用噪声表达。

9.4 代码走读:不变量清单——用事故写成的军规

AGENTS.md 的 "Invariants — do not break these" 一节,每条背后都是一次翻车:

- **Obs normalization is ON** → the normalizer must be baked into the ONNX.
  `scripts/export.py` does this; in-sim play hides the bug (it applies the
  normalizer anyway), so never hand-convert a checkpoint.
- **Policies are UNFILTERED** (no action low-pass in training). Don't add EMA
  filtering without a matched runtime flag and a transfer test — trained-with /
  deployed-without (either direction) breaks transfer.
- **Domain randomization must not accumulate across resets.** mjlab 1.3.0's
  `dr.*` ops with `operation="add"/"scale"` are natively non-accumulating (they
  re-read compile-time defaults); custom DR functions must restore-then-apply.
  An accumulating CoM randomizer once degraded every long run for months.
- If an obs is remapped to a sensor view (backlash encoder, bias), any tracking
  REWARD on the same quantity must measure the same view — otherwise the policy
  is punished for correcting what it sees.

四条军规的通用教训:

  1. 会自动修复 bug 的工具最危险:仿真回放(play)自带归一化,于是"手转 checkpoint 忘了归一化器"的 bug 在仿真里隐形、上机才爆——所以唯一的导出路径必须强制(第 10 章)。
  2. 滤波是双侧契约:训练加/部署不加(或反之)都破坏迁移——第 4 章 Tuning 注释的另一半。
  3. 随机化必须可复位:累积式 CoM 随机化"曾让每次长训劣化数月"——域随机化的 bug 不在单步,在跨 episode 的漂移,regression 测试要跨 reset 检查。
  4. 观测与奖励必须同视图:观测重映射到齿隙编码器视角后,追踪奖励也得量同一视角——否则"策略因纠正自己看见的东西而被罚"。

9.5 工作流守则:训练之前先验证物理

AGENTS.md 的 "Building a new env" 工作流里,最省钱的一条排在第二位:

2. **Verify physics assumptions in sim BEFORE training** — this is the single
   biggest time-saver:
   - A target/rest pose must be a stable equilibrium: hold its ctrl for 3 s
     from noisy inits and check TILT, not just height (a settle test that only
     records z reports fallen states as "resting fine").
   - Measure target heights off the actual robot in sim (e.g. trunk z under a
     standing policy), never carry them across model revisions. A 5 mm-wrong
     STAND_Z once turned the goal into an impossible target for days.

两个具体的坑:settle 测试只记录高度 z,会把"已经摔倒躺着"误报为"静置良好"——躺着的鸭子高度也挺稳定,必须查倾角;STAND_Z 差 5 毫米曾让目标变成物理上不可能到达的位置,"白耗数天"。训练前的物理验证是单一最大的时间节省——奖励调不通时,先怀疑目标本身不可达。

9.6 验证阶梯(修订版)

结合两个仓库的真实工具链,从训练到上机的阶梯是:

GPU 训练(mjlab,BAM 执行器)
  → play 回放(wandb run 复查指标:air_time、熵、追踪得分)
  → CPU infer_policy.py(独立运行时彩排,BAM M6 与训练一致)
  → duck-sim 软件在环(真守护进程 + MuJoCo 身体,验证部署链路)
  → 真机(保守指令起步)

每一级排除一类问题:回放排除"训练代码自身 bug",CPU 彩排排除"训练框架依赖",duck-sim 排除"守护进程与协议 bug",真机首跑只承担物理世界的残余差距。RL 仓库 scripts/ 里的 sim2real 对比脚本则专门量化"同一条策略两侧轨迹差多少"。

9.7 本章小结

Microduck 的 sim2real 三板斧,每一斧都有代码为证:执行器当被控对象建模(BAM + 摩擦随机化开在正确的口子上 + 编码器位置精确到齿隙哪一侧)、结构化误差显式建模(被动铰链 + 命名即排除 + 按步长定制接触参数)、契约用机制守护(唯一导出路径、非累积随机化、观测-奖励同视图)。外加一条元纪律:训练之前先验证物理。下一章走完最后一环:ONNX 导出与 Hub 分发。