第 8 章 microduck_rl 解剖:观测布局与奖励配方
本章代码走读:
tasks/symmetry.py(观测布局权威文档)、tasks/microduck_velocity_env_cfg.py(主配方)、tasks/mdp.py(函数库巡礼)。
8.1 代码走读:symmetry.py——61 维的"第二份真理"
第 1 章读过部署侧(obs.rs)的 61 维布局表;训练侧的权威版本写在 symmetry.py 的模块注释里——这个文件的本职是实现左右对称数据增强,注释顺手把布局和关节序固化成了文档:
Actor observation layout (61-dim flat tensor, concatenated in term insertion order):
[0:3] base_ang_vel (roll, pitch, yaw — body-frame IMU)
[3:6] projected_gravity (gx, gy, gz — body-frame)
[6:20] joint_pos_rel (14 joints, relative to default pose)
[20:34] joint_vel_rel (14 joints)
[34:48] last_action (14 joints)
[48:51] twist command (lin_vel_x, lin_vel_y, ang_vel_z)
[51:55] head command (neck_pitch, head_pitch, head_yaw, head_roll deltas)
[55:61] body command (x, y, z, roll, pitch, yaw deltas)
Joint ordering within each 14-dim block (from robot_walk.xml body tree):
0: left_hip_yaw 5: neck_pitch 9: right_hip_yaw
1: left_hip_roll 6: head_pitch 10: right_hip_roll
2: left_hip_pitch 7: head_yaw 11: right_hip_pitch
3: left_knee 8: head_roll 12: right_knee
4: left_ankle 13: right_ankle
关节序是左腿 0–4、头 5–8、右腿 9–13 的交错布局——RL 仓库的不变量清单专门警告:在 roller/backlash 模型上被动关节还会再交错,"永远不要在 mdp 函数里硬编码关节索引",要用 _servo_joint_ids 这类帮手函数。
对称增强本身是精巧的数据工程:左右镜像时,腿交换后 hip_yaw/hip_roll 取负(偏航/侧滚轴在镜像下反向),hip_pitch/knee/ankle 也要取负——不是几何原因,而是家姿态本身左右用了相反符号约定(left_hip_pitch = +0.6, right_hip_pitch = -0.6)。连"哪些维度取负"这样看似几何的问题,答案都取决于训练约定的细节。
8.2 代码走读:不变量——"零填充,永不删槽"
61 维布局能支撑策略热切换(第 4 章),靠的是一条铁律,AGENTS.md 原文:
- **Obs layout is 61D (actor) and shared across the whole policy family** so
policies are hot-swappable in the runtime: 48 base proprioception +
13D command block `[twist(3), head_pose(4), body_pose(6)]`, in that order.
An env that doesn't use a command slot ZERO-PADS it (keep the obs term,
sample tiny ranges) — never delete a slot.
不用的命令槽零填充,绝不删除。这样行走策略(用 twist)、坐站策略(用姿态标志)、拾取策略(用相位编码)共享同一个输入接口,运行时才能即插即换。第 4 章看到的"ground-pick 相位编进 twist 槽、坐站标志骑在 vx 槽"正是这条不变量在部署侧的镜像——61 维是全部策略族的公共插座。
8.3 代码走读:奖励配方——每一项权重都是一次实验
velocity 环境的奖励组装段是全书信息密度最高的代码。逐项走读:
姿态项排除头颈关节(正则的教训):
# Pose reward operates on LEG joints only. Head/neck are command-driven
# (head_pose_tracking) — if they were in this reward too, it would pull
# them to HOME while head_pose_tracking pulls them to the command, and the
# policy converges to "ignore the command" because pose reward dominates
# once head_pose_tracking's gradient dies at large commands.
cfg.rewards["pose"].params["asset_cfg"] = SceneEntityCfg(
"robot", joint_names=(r"^(?!passive_|.*neck.*|.*head.*).*",)
)
cfg.rewards["pose"].weight = 1.0
两个奖励项争夺同一组关节(一个拉向家姿态、一个拉向命令),策略的理性解是"听大头、忽略命令"——奖励冲突不是叠加,是抵消。解法用一条负向前瞻正则把头颈从 pose 项里剔除。
直立项的定量论证:
# upright: deliberately strong (2.0 / std²=0.05, was 1.0 / std²=0.1).
# 2026-07 pitch-vs-speed eval: the policy walks with a +2-4° steady forward
# lean (p90 ~6-8°) and ~2/3 of push-induced falls at speed are FORWARD. At
# weight 1.0 / std²=0.1 a 4° lean cost ~0.05/step — effectively free. At
# 2.0 / std²=0.05 it costs ~0.19/step: enough gradient to hold the trunk
# level in steady gait while transient lean (push recovery, accel) stays
# affordable.
cfg.rewards["upright"].weight = 2.0
cfg.rewards["upright"].params["std"] = math.sqrt(0.05)
一段完整的"测量 → 定价 → 调价":评估发现策略带着 +2–4° 前倾走路、2/3 的推撞跌倒是向前跌;旧参数下前倾 4° 每步只罚 0.05——"等于免费";加倍收紧后罚 0.19/步,"足够让躯干在稳态步态里保持水平,而瞬态前倾(推撞恢复、加速)仍可负担"。奖励权重不是超参数,是价格体系,调权前先算清每个行为现价多少。
故意留弱的足底打滑项:
# foot_slip deliberately weak (-0.1, not -1.0): -1.0 was too restrictive
# for this robot's pivot-heavy turning.
cfg.rewards["foot_slip"].weight = -0.1
这只鸭子转弯靠 pivot(原地碾转),足底打滑是转弯方式的一部分,罚重了转不了弯。
原地转的采样修补:
# Fraction of envs commanded to spin on the spot (lin=0, |ang| ∈ [0.4·max, max]).
TURN_IN_PLACE_FRACTION = 0.15
头注释交代了病因:"independent uniform sampling makes spin-on-the-spot ~2% of data → untrained"(2026-07 审计)——速度与角速度独立均匀采样时,"线速度为零、角速度大"的原地转只占数据 2%,策略压根没学过。修法不是改奖励,是改指令分布:15% 的环境强制分派原地转命令。RL 调参的一半功夫在检查"策略到底见过什么"。
8.4 代码走读:头部下垂之战——一次教科书级的奖励手术
2026 年 8 月 20 日的修复记录,是整个仓库最精彩的一段注释,完整还原了一次失败的直接修法与成功的迂回修法:
# Head droop fix (2026-08-20). The head walks pitched ~15° down (measured:
# run ww1g2198 head_pose_tracking 1.544/2.0 → 14.6° mean joint error).
# DO NOT fix this by tightening head_pose_tracking's std: run 5yay13u4 tried
# fine_std=0.1 and the policy stopped walking entirely by iter 300 (air_time
# 1.01 → 0.02, peak foot height 15 mm → 2 mm, entropy collapsed 10.9 → 1.9).
# An instantaneous tight tolerance taxes walking 0.77/step — 76% of the whole
# air_time reward — and is UNESCAPABLE, since a 280 g head (38% of robot
# mass) must oscillate while stepping. Standing still scored higher, so it
# stood still.
# The DC bias, unlike the oscillation, IS escapable (bias the neck command up
# to cancel gravity sag), so price only that: L1 on a 1 s EMA of the error.
# At the optimum this costs a walking policy nothing.
cfg.rewards["head_pose_bias"] = RewardTermCfg(
func=microduck_mdp.head_pose_bias_penalty,
weight=0.0, # ramped by the head_pose_bias_weight curriculum below
params={"command_name": "_head_pose", "tau_s": 1.0},
)
逐句拆解这个案例:
- 症状量化:走路时头前俯约 15°;run
ww1g2198的头部追踪得分 1.544/2.0,折合平均关节误差 14.6°。注意他们用 wandb run id + 指标数字记录问题——可复现是调试的前提。 - 直接修法及其死亡:收紧
head_pose_tracking的容差(std=0.1),run5yay13u4——策略 300 迭代后彻底不走了:空中时间 1.01 → 0.02,峰值抬脚高度 15mm → 2mm,策略熵 10.9 → 1.9(坍缩)。 - 死因分析:瞬时紧容差对行走课税 0.77/步——占全部空中时间奖励的 76%——且这笔税"不可逃避",因为 280 g 的头占整机质量 38%,迈步时必然振荡,瞬时误差永远存在。站着不动得分反而更高,于是策略学会站着不动。"DO NOT fix this by tightening std"——开头这句禁令就是给未来的人省下 retrain 一次的学费。
- 迂回修法:区分可逃避与不可逃避的误差——振荡不可逃避(不罚),直流偏差可逃避(抬头颈命令抵消重力下垂即可消除),于是只对误差的 1 秒指数滑动平均(EMA,τ=1s)收 L1 罚。最优点上,走得好且头不垂的策略付税为零。
这一段的教学价值超过十篇论文:奖励设计的对象不是"错误",是"可逃避的错误"。对不可逃避的误差收税,等于教策略放弃整个任务。
8.5 mdp.py 函数库巡礼
7400 行的 mdp.py 收拢约 80 个函数,函数名即奖励配方全景(节选分组):
- 跌倒与恢复:
_fallen_mask、feet_air_time_upright、upright_progress、height_progress、fallen_state_penalty、recovery_success、fallen_too_long - 执行器健康:
servo_stall_penalty(舵机堵转)、servo_acc_spike_penalty(加速度尖峰) - 动作平滑(带跌倒豁免):
action_rate_l2_fallen_scaled、joint_torque_rate_l2_fallen_scaled——名字里的 "fallen_scaled" 意味着跌倒状态下放松动作惩罚:都已经摔了,别再为甩腿的猛烈动作扣分,先学爬起来 - 分部位平滑:
leg_action_rate_l2/neck_action_rate_l2分开——腿和脖子的"猛烈"标准不同 - 数值防御:
_nan_safe_reward_compute(NaN 安全的奖励计算)、robot_state_is_nan(NaN 终止)、_safe_compute_returns - 重置技巧:
reset_with_forward_velocity(带初速重置,练加速稳定性)、reset_action_history、reset_rolling_entry - 轮滑专用:
wheel_glide_reward、crouch_glide_*系列
两个系统级细节:_nan_safe_reward_compute 的存在说明 NaN 不是异常而是要常态化防御的工况(GPU 并行仿真里一个环境发散不能烧掉整批梯度);分部位、分状态的惩罚变体说明"平滑"从来不是一个项,是一族项。
8.6 本章小结
奖励配方的三条心法,全部来自上面的代码:权重是价格,调权先算现价(upright);奖励冲突会抵消,同组关节只能有一个主人(pose 正则);只罚可逃避的错误(head_pose_bias 的 EMA 手术)。观测侧的心法只有一条但最硬:61 维是全策略族的公共插座,零填充、永不删槽。下一章把这些配方放进 sim2real 的战场。