第 8 章 microduck_rl 解剖:观测布局与奖励配方

本章代码走读:tasks/symmetry.py(观测布局权威文档)、tasks/microduck_velocity_env_cfg.py(主配方)、tasks/mdp.py(函数库巡礼)。

8.1 代码走读:symmetry.py——61 维的"第二份真理"

第 1 章读过部署侧(obs.rs)的 61 维布局表;训练侧的权威版本写在 symmetry.py 的模块注释里——这个文件的本职是实现左右对称数据增强,注释顺手把布局和关节序固化成了文档:

Actor observation layout (61-dim flat tensor, concatenated in term insertion order):
    [0:3]   base_ang_vel      (roll, pitch, yaw  — body-frame IMU)
    [3:6]   projected_gravity (gx, gy, gz         — body-frame)
    [6:20]  joint_pos_rel     (14 joints, relative to default pose)
    [20:34] joint_vel_rel     (14 joints)
    [34:48] last_action       (14 joints)
    [48:51] twist command     (lin_vel_x, lin_vel_y, ang_vel_z)
    [51:55] head command      (neck_pitch, head_pitch, head_yaw, head_roll deltas)
    [55:61] body command      (x, y, z, roll, pitch, yaw deltas)

Joint ordering within each 14-dim block (from robot_walk.xml body tree):
    0: left_hip_yaw    5: neck_pitch    9:  right_hip_yaw
    1: left_hip_roll   6: head_pitch    10: right_hip_roll
    2: left_hip_pitch  7: head_yaw      11: right_hip_pitch
    3: left_knee       8: head_roll     12: right_knee
    4: left_ankle                       13: right_ankle

关节序是左腿 0–4、头 5–8、右腿 9–13 的交错布局——RL 仓库的不变量清单专门警告:在 roller/backlash 模型上被动关节还会再交错,"永远不要在 mdp 函数里硬编码关节索引",要用 _servo_joint_ids 这类帮手函数。

对称增强本身是精巧的数据工程:左右镜像时,腿交换后 hip_yaw/hip_roll 取负(偏航/侧滚轴在镜像下反向),hip_pitch/knee/ankle 也要取负——不是几何原因,而是家姿态本身左右用了相反符号约定(left_hip_pitch = +0.6, right_hip_pitch = -0.6)。连"哪些维度取负"这样看似几何的问题,答案都取决于训练约定的细节。

8.2 代码走读:不变量——"零填充,永不删槽"

61 维布局能支撑策略热切换(第 4 章),靠的是一条铁律,AGENTS.md 原文:

- **Obs layout is 61D (actor) and shared across the whole policy family** so
  policies are hot-swappable in the runtime: 48 base proprioception +
  13D command block `[twist(3), head_pose(4), body_pose(6)]`, in that order.
  An env that doesn't use a command slot ZERO-PADS it (keep the obs term,
  sample tiny ranges) — never delete a slot.

不用的命令槽零填充,绝不删除。这样行走策略(用 twist)、坐站策略(用姿态标志)、拾取策略(用相位编码)共享同一个输入接口,运行时才能即插即换。第 4 章看到的"ground-pick 相位编进 twist 槽、坐站标志骑在 vx 槽"正是这条不变量在部署侧的镜像——61 维是全部策略族的公共插座。

8.3 代码走读:奖励配方——每一项权重都是一次实验

velocity 环境的奖励组装段是全书信息密度最高的代码。逐项走读:

姿态项排除头颈关节(正则的教训):

    # Pose reward operates on LEG joints only. Head/neck are command-driven
    # (head_pose_tracking) — if they were in this reward too, it would pull
    # them to HOME while head_pose_tracking pulls them to the command, and the
    # policy converges to "ignore the command" because pose reward dominates
    # once head_pose_tracking's gradient dies at large commands.
    cfg.rewards["pose"].params["asset_cfg"] = SceneEntityCfg(
        "robot", joint_names=(r"^(?!passive_|.*neck.*|.*head.*).*",)
    )
    cfg.rewards["pose"].weight = 1.0

两个奖励项争夺同一组关节(一个拉向家姿态、一个拉向命令),策略的理性解是"听大头、忽略命令"——奖励冲突不是叠加,是抵消。解法用一条负向前瞻正则把头颈从 pose 项里剔除。

直立项的定量论证:

    # upright: deliberately strong (2.0 / std²=0.05, was 1.0 / std²=0.1).
    # 2026-07 pitch-vs-speed eval: the policy walks with a +2-4° steady forward
    # lean (p90 ~6-8°) and ~2/3 of push-induced falls at speed are FORWARD. At
    # weight 1.0 / std²=0.1 a 4° lean cost ~0.05/step — effectively free. At
    # 2.0 / std²=0.05 it costs ~0.19/step: enough gradient to hold the trunk
    # level in steady gait while transient lean (push recovery, accel) stays
    # affordable.
    cfg.rewards["upright"].weight = 2.0
    cfg.rewards["upright"].params["std"] = math.sqrt(0.05)

一段完整的"测量 → 定价 → 调价":评估发现策略带着 +2–4° 前倾走路、2/3 的推撞跌倒是向前跌;旧参数下前倾 4° 每步只罚 0.05——"等于免费";加倍收紧后罚 0.19/步,"足够让躯干在稳态步态里保持水平,而瞬态前倾(推撞恢复、加速)仍可负担"。奖励权重不是超参数,是价格体系,调权前先算清每个行为现价多少。

故意留弱的足底打滑项:

    # foot_slip deliberately weak (-0.1, not -1.0): -1.0 was too restrictive
    # for this robot's pivot-heavy turning.
    cfg.rewards["foot_slip"].weight = -0.1

这只鸭子转弯靠 pivot(原地碾转),足底打滑是转弯方式的一部分,罚重了转不了弯。

原地转的采样修补:

# Fraction of envs commanded to spin on the spot (lin=0, |ang| ∈ [0.4·max, max]).
TURN_IN_PLACE_FRACTION = 0.15

头注释交代了病因:"independent uniform sampling makes spin-on-the-spot ~2% of data → untrained"(2026-07 审计)——速度与角速度独立均匀采样时,"线速度为零、角速度大"的原地转只占数据 2%,策略压根没学过。修法不是改奖励,是改指令分布:15% 的环境强制分派原地转命令。RL 调参的一半功夫在检查"策略到底见过什么"。

8.4 代码走读:头部下垂之战——一次教科书级的奖励手术

2026 年 8 月 20 日的修复记录,是整个仓库最精彩的一段注释,完整还原了一次失败的直接修法与成功的迂回修法:

    # Head droop fix (2026-08-20). The head walks pitched ~15° down (measured:
    # run ww1g2198 head_pose_tracking 1.544/2.0 → 14.6° mean joint error).
    # DO NOT fix this by tightening head_pose_tracking's std: run 5yay13u4 tried
    # fine_std=0.1 and the policy stopped walking entirely by iter 300 (air_time
    # 1.01 → 0.02, peak foot height 15 mm → 2 mm, entropy collapsed 10.9 → 1.9).
    # An instantaneous tight tolerance taxes walking 0.77/step — 76% of the whole
    # air_time reward — and is UNESCAPABLE, since a 280 g head (38% of robot
    # mass) must oscillate while stepping. Standing still scored higher, so it
    # stood still.
    # The DC bias, unlike the oscillation, IS escapable (bias the neck command up
    # to cancel gravity sag), so price only that: L1 on a 1 s EMA of the error.
    # At the optimum this costs a walking policy nothing.
    cfg.rewards["head_pose_bias"] = RewardTermCfg(
        func=microduck_mdp.head_pose_bias_penalty,
        weight=0.0,  # ramped by the head_pose_bias_weight curriculum below
        params={"command_name": "_head_pose", "tau_s": 1.0},
    )

逐句拆解这个案例:

  1. 症状量化:走路时头前俯约 15°;run ww1g2198 的头部追踪得分 1.544/2.0,折合平均关节误差 14.6°。注意他们用 wandb run id + 指标数字记录问题——可复现是调试的前提。
  2. 直接修法及其死亡:收紧 head_pose_tracking 的容差(std=0.1),run 5yay13u4——策略 300 迭代后彻底不走了:空中时间 1.01 → 0.02,峰值抬脚高度 15mm → 2mm,策略熵 10.9 → 1.9(坍缩)。
  3. 死因分析:瞬时紧容差对行走课税 0.77/步——占全部空中时间奖励的 76%——且这笔税"不可逃避",因为 280 g 的头占整机质量 38%,迈步时必然振荡,瞬时误差永远存在。站着不动得分反而更高,于是策略学会站着不动。"DO NOT fix this by tightening std"——开头这句禁令就是给未来的人省下 retrain 一次的学费。
  4. 迂回修法:区分可逃避与不可逃避的误差——振荡不可逃避(不罚),直流偏差可逃避(抬头颈命令抵消重力下垂即可消除),于是只对误差的 1 秒指数滑动平均(EMA,τ=1s)收 L1 罚。最优点上,走得好且头不垂的策略付税为零。

这一段的教学价值超过十篇论文:奖励设计的对象不是"错误",是"可逃避的错误"。对不可逃避的误差收税,等于教策略放弃整个任务。

8.5 mdp.py 函数库巡礼

7400 行的 mdp.py 收拢约 80 个函数,函数名即奖励配方全景(节选分组):

两个系统级细节:_nan_safe_reward_compute 的存在说明 NaN 不是异常而是要常态化防御的工况(GPU 并行仿真里一个环境发散不能烧掉整批梯度);分部位、分状态的惩罚变体说明"平滑"从来不是一个项,是一族项。

8.6 本章小结

奖励配方的三条心法,全部来自上面的代码:权重是价格,调权先算现价(upright);奖励冲突会抵消,同组关节只能有一个主人(pose 正则);只罚可逃避的错误(head_pose_bias 的 EMA 手术)。观测侧的心法只有一条但最硬:61 维是全策略族的公共插座,零填充、永不删槽。下一章把这些配方放进 sim2real 的战场。