第 4 章 robotd:50 Hz 控制循环与技能调度
本章代码走读:
robotd/src/control.rs(约 750 行,纯计算核心)、robotd/src/main.rs(约 8100 行,IO 与常量)、duck-control/src/obs.rs(观测组装)。
4.1 一图流:一个 tick 里发生什么
control.rs 的模块注释是全仓库最好的一份"程序说明书",值得整段精读:
//! Turning sensors and a command into joint targets — and scheduling the skills.
//!
//! Everything here is pure computation between [`duck_control::io::RobotIo::read`] and the
//! safety layer's `apply`. It holds no IO handle — by construction it cannot command a
//! motor, only propose targets.
//!
//! The tick, in order:
//!
//! ```text
//! skill windows ← advance / expire (roulade window, kick timer, ground-pick phase, sit↔stand rise)
//! command ← the caller's smoothed command, re-encoded for the active skill
//! net ← roulade > kick > ground pick > sit/rise > stand-by-magnitude > walk
//! action ← ONNX
//! targets ← home pose + action_scale × action
//! filters ← optional first-order low-pass on head and legs
//! ```
第一段划定了可测试性边界:控制器不持有任何 IO 句柄,"按构造它无法命令电机,只能提议目标"。读传感器 → 纯计算 → 安全层执行,三段切开后,中间那段可以在任何机器上做单元测试(第 6 章的仿真也受益于此)。
末行公式是策略输出的完整语义:关节目标 = 家姿态 + action_scale × 策略输出。策略学的不是绝对角度,而是围绕家姿态的偏移——这个约定必须训练与部署两侧一致。
4.2 代码走读:Tuning——每个数字都对着训练侧
pub struct Tuning {
/// Scales raw policy output before it becomes a joint offset. The prototype's current
/// alpha default.
pub action_scale: f64,
/// The standing policy is trained to be applied whole.
pub standing_action_scale: f64,
/// Standing runs softer, at this fraction of the running gain. `--standing-kp-ratio`.
pub standing_gain_ratio: f64,
pub gain: u16,
/// First-order low-pass on the head joints. `None` is no filtering. The alpha policies
/// are trained with 0.5 — it must match training or transfer degrades.
pub head_lowpass: Option<f64>,
/// Same, for the ten leg joints. Trained with 0.7.
pub legs_lowpass: Option<f64>,
}
impl Default for Tuning {
fn default() -> Self {
Self {
action_scale: 0.9,
standing_action_scale: 1.0,
standing_gain_ratio: 0.8,
gain: 200,
head_lowpass: Some(0.5),
legs_lowpass: Some(0.7),
}
}
}
注意 head_lowpass: Some(0.5) 的注释——"alpha 策略是以 0.5 训练的,必须与训练一致,否则迁移退化"。低通滤波不是部署侧可以自由调节的"手感参数",而是训练契约的一部分(RL 仓库 AGENTS.md 从另一侧说了同一件事:策略默认无滤波,单方面在任一侧加/改滤波都会破坏迁移)。两个仓库在这类细节上的互文,就是 sim2real 的"契约面"。
4.3 代码走读:网络选择优先级链与相位编码
step() 的后半段是全章最密的逻辑——根据当前状态选择驱动网络,并把状态编码进命令槽:
} else if let Some(phase) = self.ground_pick {
// The twist slots carry the phase encoding; head and body are zero-padded,
// mirroring the training env's `zero_command_padding`.
let angle = std::f64::consts::TAU * phase;
let c = Command {
twist: [angle.cos(), angle.sin(), 0.0],
..Command::default()
};
(Net::GroundPick, c, "ground_pick".into())
} else {
let mut c = *command;
match self.sit {
// The posture flag rides the twist vx slot: 1 = sit, 0 = stand. Head and
// body slots stay live — the prototype keeps them in the buffer too.
Sit::Sitting => {
c.twist = [1.0, 0.0, 0.0];
(Net::SitStand, c, "sit".into())
}
两处"盗用"命令槽的编码:
- ground-pick 的进度相位被编成
(cos 2πφ, sin 2πφ)塞进 twist 的 vx/vy 槽——用两个正交分量表达一个循环相位,策略能从观测里无歧义地解码"拾取动作进行到哪一步",且 φ=0 与 φ=1 在圆上是同一点(天然周期连续)。 - 坐/站姿态标志骑在 twist vx 槽上:1 = 坐,0 = 站。注释特意说明头和身体槽保持活跃。
这类"一槽多用"是对 61 维统一观测的极致利用:不改布局,就能让不同任务族在同一个向量里携带各自需要的信号。
优先级链本身(模块注释里的 roulade > kick > ground pick > sit/rise > stand-by-magnitude > walk)在代码中体现为 if-else 的顺序;注释还保留了两个容易"好心修坏"的遗产:
//! - **A kick window runs at standing tuning.** The kick's observation carries an all-zero
//! command, and in the prototype the standing transition fires on exactly that — so a
//! kick runs at `standing_action_scale` and the softened standing gain. Kept, because
//! the kicks were tuned against it.
踢球窗口的观测命令全零,恰好触发了"切站立调参"的分支——踢球实际跑在站立的柔和增益下。**这是缺陷还是特性?团队的选择:保留,因为踢球就是在这个条件下调出来的。**对行为系统而言,"历史一致性"有时比"逻辑纯洁性"更值钱。
4.4 代码走读:COAST_TICKS——一次"随机微小抽搐"事故的尸检
main.rs 里最值得全文摘录的注释,讲的是如何对待例行故障:
/// **One dropped Dynamixel transaction is ordinary** — `robotd` says so itself when it logs
/// one — and it must therefore be *invisible*. It was not: a failed read used to stop the
/// policy for that tick, which commanded the hold pose and reset the controller, so every
// ordinary dropped read produced a visible twitch. Measured on a bench robot at ~8 drops a
/// minute with a monitor attached, which is exactly the reported "random tiny spasms".
///
/// Coasting is safe because the observation is *already* a tick old by construction: at
/// 50 Hz the policy is trained on data of exactly this age, and a second tick of it is
/// inside that. Three ticks (60 ms) covers a drop, a retry and a slow tick; past that the
/// robot genuinely cannot see, and holding still is the honest answer.
const COAST_TICKS: u32 = 3;
完整的问题-诊断-论证-参数选择链条:
- 现象:用户报告"随机微小抽搐";台架实测每分钟约 8 次总线丢包。
- 病因:一次读失败就停策略一拍 → 指令保持姿态并复位控制器 → 每次例行丢包都变成一次可见抽动。
- 修法的合理性证明:观测本来就滞后一拍(50 Hz 下策略就是按这个数据年龄训练的),惯性滑行一拍在训练分布之内;三拍(60 ms)覆盖"一次丢包 + 一次重试 + 一次慢拍",超过三拍才真的看不见,"保持静止才是诚实的答案"。
配套的 RESET_AFTER_PAUSE(200 ms)同理:控制器复位本身是不连续动作,20 ms 的网络抖动不配触发它,200 ms 的真中断才配。故障处理的粒度要和故障的真实频率对齐——这是嵌入式控制系统的通用课。
4.5 代码走读:观测组装的两条禁令
第 1 章看过 61 维布局表,obs.rs 随后列出了"容易做错、且是对着原型逐项确认过而非想当然"的两条:
//! 1. **Body x, y and yaw are hardcoded zero.** They are unbound in the training
//! environment, so an all-zero body command is the *nominal* encoding, not a
//! placeholder standing in for something better.
//! 2. **Head targets ride in the command, and are not added on top of the policy output.**
//! The prototype does both, in different modes, and gates the post-hoc addition behind
//! `if !new_cmd_obs` with the note "head_offsets are a COMMAND fed via the obs vector
//! instead — don't double-add it here". Doing both would bend the head twice.
第二条尤其典型:头部目标是喂给策略的命令,不能在策略输出之后再叠加一次——"两处都做会把头掰弯两次"。跨系统接口的参数,"谁消费"必须唯一。
4.6 循环与 IPC 的关系
main.rs 的开篇注释交代了 8100 行 daemon 与控制循环的边界:IPC 侧"publishes and never calls into the loop"——只发布、绝不回调进循环;一个卡死的循环靠自己报告不健康("a wedged loop reports itself unhealthy"),而不是被外部探测杀死。日志预算同样体现纪律:每 tick 都记日志的话"50 Hz 一天约 430 万行",所以常态静默、每 5 分钟(LOOP_SUMMARY_INTERVAL)汇总一次,孤立丢包 60 秒内不重复抱怨。
4.7 策略侧:槽位、递归与热切换
从 Call 枚举(第 3 章)已经见过 RobotPolicies / RobotLoadPolicy / RobotReloadPolicies 与 PolicySlot 结构——robotd 按槽位管理策略(行走、站立、坐站、拾取、技能各占一槽),加载即热切换,且支持"重读全部槽位"。main.rs 注释还提到:"API 2 adds explicit recurrent state; API 1 feed-forward models remain supported"——循环策略(带隐状态)已在协议里占位,前馈模型继续兼容(详见 docs/recurrent-policies.md)。
4.8 本章小结
robotd 的工程设计可以提炼成四条:纯计算核心(可测试性按构造保证)、调参即契约(每个数字对着训练侧)、例行故障必须不可见(滑行三拍)、接口参数消费唯一(不许掰弯两次)。8100 行里最珍贵的不是算法,是这些把"上机才暴露的问题"提前写进代码纪律的注释。下一章快速巡览其余守护进程。