Microduck 开源软件剖析

从守护进程到 Sim-to-Real
v0.2 扩写版 · 每章附真实源码走读
基于 pollen-robotics/microduck 与 microduck_rl(Apache-2.0)
2026 年 9 月

目录

前言

2026 年 8 月,Pollen Robotics(已并入 Hugging Face)发布了 Microduck:一台 25 厘米高、约 800 克、售价 399 美元的双足机器鸭。它真正的价值不在硬件——14 个 Dynamixel 舵机、一颗 Rockchip SoC,这些都是成熟供应链里的东西——而在于它把一整套从仿真训练到真机部署的物理 AI 流水线连同机器人操作系统全部开源了出来。

RL 仓库 AGENTS.md 的第一段有一句话,是这本书的题眼:

Sim2real transfer is the whole point: every convention below exists because breaking it produced a policy that worked in the viewer and failed on hardware.

(Sim2real 迁移就是全部意义:下文每一条约定的存在,都是因为打破它曾产出一个在查看器里工作、上机就翻车的策略。)

v0.2 的这次扩写,做的就是把这些"约定"逐条挖出来的工作。两个仓库克隆到本地逐文件读过后,我发现它们最珍贵的不是架构图,而是注释里的事故记录:哪个 wandb run 用 0.1 的容差把策略练到再也不走路、哪次累积式随机化让长训劣化了数月、哪块板子因为两处版本漂移装上了会让运行时 panic 的旧库。这些内容教科书里没有,论文里更没有——它们只存在于"把系统真正跑起来的地方"。

因此 v0.2 每章都加了代码走读:引用一律逐字摘录(含原注释与实验数字),读者可以拿着书对着仓库读。全书的高光段落包括:

如何读这本书:有 Rust 基础的读者顺序读上半部(1–6 章)再跳第 10 章;做强化学习的读者从第 7 章直入(第 8、9 章是重点);只想跑通流程的读者读第 6、7 章加各章末的命令清单。每章开头的引言块标注了该章走读的源码文件,附录 B 有完整的文件-章节对照表。

致谢:所有素材来自 Pollen Robotics 的公开博客、两个 GitHub 仓库(Apache-2.0)及社区报道。特别感谢把失败实验写进注释的工程师们——那些注释是这本书真正的作者。

第 1 章 初识 Microduck

本章代码走读:Cargo.toml(workspace 根)、AGENTS.md(RL 仓库)、duck-control/src/obs.rs 开篇。

1.1 它是什么

Microduck 是 Pollen Robotics 推出的小型双足机器人,2026 年 8 月 27 日开启预售,定价 399 美元(税前),目标 2026 年圣诞节前在北美、欧洲和英国首发。四款配色:Cream、Graphite、Lavender、Sky。

硬件规格:

一个容易搞混的数字就此澄清:媒体报道的"15 电机"与 RL 仓库的"14 XL330"都对——策略世界是 14 维的,嘴是策略世界之外的第 15 个执行器。

1.2 软件全景与两个仓库的分工

软件以 Apache-2.0 开源在 GitHub 的 pollen-robotics 组织下(3D 模型文件为 CC BY-SA-NC):

pollen-robotics/microduck       # 机器人软件 + SDK:Rust workspace,22 个成员 crate
pollen-robotics/microduck_rl    # 强化学习训练:Python/mjlab(MuJoCo Warp)+ rsl_rl PPO

两者的接口契约只有两个文件大小:一个 61 维观测向量进,14 维动作出 的 ONNX 图,外加一份 manifest.json。RL 仓库的 AGENTS.md 第一段就把这条链路讲清楚了:

RL training environments for Microduck — a ~800 g, ~25 cm tall bipedal
robot with 14 Dynamixel XL330 servos — built on mjlab
(MuJoCo Warp) with PPO (rsl_rl). Policies are trained here at 50 Hz, exported to
ONNX, and deployed by the runtime in the `pollen-robotics/microduck` repo on
the real robot. Sim2real transfer is the whole point: every convention below
exists because breaking it produced a policy that worked in the viewer and
failed on hardware.

最后一句值得抄在工位上:"每一条约定之所以存在,都是因为打破它曾产出一个在查看器里工作、上机就翻车的策略。" 这句话是全书的题眼——后面所有章节,无论是 Rust 侧的观测组装还是 Python 侧的奖励调参,讲的都是这类"看不见的契约"。

1.3 代码走读:workspace 根 Cargo.toml

主仓库根目录的 Cargo.toml 是全书的第一个"架构文档"。它不只列成员,还把三个外部依赖的版本钉死策略写在 [workspace.metadata] 里,每一段注释都是一个真实事故的墓志铭:

# Workspace root for the robot daemon.
#
# One crate per service or tool, matching the service split in
# docs/design/architecture.md §1. `robotd`, `btd` and `mediad` are all siblings now.
[workspace]
resolver = "3"
members = ["btd", "configd", "duck-control", "duck-detect", "duck-ether", "duck-ipc-proto",
  "duckctl", "kinematics", "mediad", "odometry", "pad-imu", "padd", "pet-detect", "sounds",
  "tof", "updater", "robotctl", "robotd", "robotd-params", "test-support", "uyvy", "xtask"]

三段元数据(节选):

# ONNX Runtime, which robotd dlopens to run a policy. One source of truth: `xtask package`
# bakes these into the release's preinstall hook, and a test asserts scripts/setup-board.sh
# agrees with them. Two copies drifting apart is exactly how a board ended up with 1.20.1
# against an `ort` that panics below 1.23.
[workspace.metadata.onnxruntime]
floor  = "1.23"
target = "1.28.0"

# The official policy set, which `robotd` runs and which lives on the Hugging Face Hub rather
# than in this repository — a gait retrain should not need a daemon release, and a daemon fix
# should not re-ship six megabytes of unchanged weights.
[workspace.metadata.policies]
repo    = "pollen-robotics/microduck-policies"
version = "v5"

三个设计决策藏在这里:

  1. ONNX Runtime 是运行时 dlopen 的,不是编译期链接——所以它的版本约束必须有人守着,守的方式是"元数据 + 安装脚本 + 一个断言两者一致的测试"三方互锁;
  2. 官方策略集住在 Hugging Face Hub(microduck-policies,最低版本 v5),不进代码仓库——"步态重训不应该要求守护进程发版,守护进程修复也不应该重新分发六兆没变的权重"。这句注释把"策略即资产、与代码解耦"的部署哲学一句话讲完;
  3. Rockchip NPU 运行时(rknn v2.3.2)同样钉死——"比模型旧的运行时会在 rknn_init 处带着一个数字失败",这种注释只有踩过坑才写得出来。

1.4 代码走读:duck-control/src/obs.rs 开篇——"最高风险代码"

主仓库里 duck-control crate 的 obs.rs,模块注释开头一句就给全仓库的风险排了序:

//! The observation vector the policy sees.
//!
//! **This is the highest-risk code in the crate.** It is a flat array of 61 floats whose
//! every index must match what the policy was trained against. A wrong offset does not fail
//! loudly — it produces a plausible-looking robot that falls over, and the symptom looks
//! like a tuning or timing problem rather than an indexing one.

61 维观测的完整布局(这段表在部署侧和训练侧各有一份,必须逐位一致):

index   width  contents
0..3        3  gyro, trunk frame, rad/s
3..6        3  projected gravity, trunk frame, unit vector
6..20      14  joint position minus home pose, mouth excluded
20..34     14  joint velocity, mouth excluded
34..48     14  previous action, mouth excluded
48..61     13  command (below)

命令块 13 维:速度指令 3(vx、vy、vyaw)+ 头部姿态 4 + 身体姿态 6。第 8 章会从训练侧看到同一张表的镜像,第 4 章会看到 robotd 如何每个 tick 把它组装出来。

为什么一个 61 个浮点数的数组是"最高风险代码"? 因为它失败的形态不是报错,而是"一只看起来合理但站不住的鸭子"——症状像调参问题,病因却是索引错位。这是所有跨进程、跨语言 ML 系统的通病,Microduck 用文档加注释加双侧常量断言来防御(duck_ipc_proto::POLICY_OBS_LEN 与本 crate 的 OBS_LEN 有编译期一致性断言)。

1.5 为什么它值得研究

  1. 小而完整,且注释即历史。两个仓库的代码注释里大量记录了"哪次实验、哪个 run、什么数字、为什么改"——不是教学示例代码的整洁,而是生产系统的诚实。
  2. sim2real 的全链路样板。从 Onshape CAD 导出 MJCF、BAM 执行器建模、齿隙注入、域随机化、ONNX 导出烘焙归一化器、Hub 分发、守护进程加载——每一环都有真实代码。
  3. 消费级产品工程。配网、蓝牙配对、WebRTC、签名更新、日志预算——研究原型最欠的债这里都还了。

1.6 本章小结

Microduck = 14 个策略舵机 + 1 个嘴 + 一颗 RK3566 + 两个开源仓库。两个仓库之间只隔着一张 [1,61] → [1,14] 的 ONNX 图,而这个接口的每一侧都把"对齐"当作最高风险事项对待。下一章进入 Rust workspace 的内部结构。

第 2 章 总体架构:Rust workspace 与守护进程群

本章代码走读:Cargo.toml 的 default-members、AGENTS.md(主仓库)、docs/ 目录。

2.1 22 个 crate 的分工

主仓库是一个单 workspace,成员一眼可以分成四类:

类别 crate 说明
守护进程 robotd updater configd btd padd mediad tof 机器人上常驻,一服务一进程
领域库 duck-control duck-ipc-proto duck-detect duck-ether kinematics odometry pad-imu pet-detect sounds robotd-params uyvy 纯逻辑或纯协议,被守护进程与工具复用
客户端工具 robotctl duckctl 机上 CLI / 开发者笔记本 CLI
构建支撑 test-support xtask deploy(脚本) 测试夹具、自定义 cargo 子命令、部署

体量分布很说明问题:最大的单文件是 robotd/src/main.rs(约 8100 行,第 4 章主角)和 duck-ipc-proto/src/lib.rs(约 6200 行,第 3 章主角)——系统的心脏是"控制循环"和"契约"这两块,其余 crate 都小得多。控制器的核心计算被抽到 robotd/src/control.rs(约 750 行)保持可读,这是"大 daemon、小核心"的典型布局。

注意 tof(深度传感器)——博客里说的 tofd 守护进程对应的 crate 名是 tof;同样,更新服务 crate 名是 updater(daemon 二进制名 updaterd)。命名上"d"后缀标记 daemon 二进制,crate 名不带。

2.2 代码走读:default-members 与一台不该看到蓝牙栈的开发机

第 1 章读过 members,紧随其后的 default-members 有一段仓库里最生动的注释——duckctl(笔记本上的客户端)被排除在板机构建之外的原因:

# Everything except `duckctl`, and that exception is the whole reason this key exists.
#
# `cargo board --bins` — in `dev-push.sh` and in the release workflow — builds every default
# member for aarch64, which is right for a robot's daemons and wrong for a client a developer
# runs on their own machine: it would cross-compile a Bluetooth stack for a board that must never
# see one, on the release path, for nothing.
#
# The alternative was naming the binaries explicitly at each `--bins` call site. There are two of
# them and they would have to be kept in step by hand, which is the failure this repo keeps
# writing down. One list here, and a new daemon is picked up by both without anybody remembering.

一层意思浮在表面(省一次无用的交叉编译),另一层是方法论:"两处手工保持同步"被这个仓库视为必须用机制消灭的故障模式——要么收进一个列表,要么写个测试盯着。这个思想在后文反复出现:ONNX Runtime 版本三方互锁、PIN 路由表测试、策略 manifest 双侧校验,全是同一招。

2.3 文档体系:机制唯一归属

docs/ 的组织在这个仓库被上升到制度。主仓库 AGENTS.md 开篇立规矩:

## Docs own mechanisms; one page each

`docs/README.md` assigns every mechanism to one design doc. When a fact belongs to a page listed
there, every other page says one sentence and links. When two pages disagree, the one that does
not own the mechanism is the bug — and when behaviour and a design doc disagree, the doc is
the bug.

设计文档目录 docs/design/ 现有 11 篇,篇名即机制清单:

app-path-design.md     architecture.md        boot-recovery-net.md
policy-channel-design.md  remote-access-design.md  remote-webrtc.md
restart-order.md       robotd-design.md       simulation.md
updater-design.md      webrtc-console.md

另有 policy-manifest.md(策略清单契约——第 10 章的主角,且 RL 仓库的发布代码显式引用它)、recurrent-policies.md(循环策略 API)、faq.md(面向"对着一只鸭子开发"的人的任务型入口)。

对写书人的启示:这个仓库里注释与文档不是代码的附庸,而是决策记录。当行为与文档不一致时,被判死刑的是文档——因为腐化的文档比没有文档更危险。

2.4 三条产品级守则

主仓库 AGENTS.md 还有三条容易被研究型团队忽视的守则,值得全文摘录:

其一,消费级视频链路的主备关系:

## A consumer uses WebRTC. `media.stream` is the fallback

The robot publishes H.264 over WebRTC, and that is the default for anything consuming a duck's
camera — it is encrypted end to end, it carries the control channel on the same session, and it
has a return path. ...

This is worth stating because the repository reads the other way round if you only follow the
code: `media.stream` was built when the relay endpoint was dead and WebRTC genuinely could not
connect from a data centre, so its module doc argues its own case at length. That endpoint is
fixed. Do not conclude from the volume of prose that it is the preferred path.

——"不要因为某个模块的注释嗓门大就以为它是主路径"。代码库会留下历史的沉积层,读库要区分"现在时"与"过去时"。

其二,永不围绕版本差异做设计:老客户端的局限是拿来提版本号的理由,不是绕行的理由;版本偏差被记录并继续服务,只有真正缺失的路由或未知参数才可以拒绝。

其三,release 才是修复到达机器人的方式:"main being fixed is not a robot being fixed."——机器人走 stable 通道,修复要切 release 才到用户手里。研究代码没有这个概念,产品代码必须有。

2.5 守护进程模型回顾

对照第 1 章的表:robotd(控制平面)、updaterd(维护平面)、configd/btd/padd(连接与输入平面)、mediad/tofd(媒体与感知平面)。故障隔离、权限分离、独立演进的好处上一版已述;代价——IPC——是下一章整章的主题。

2.6 本章小结

架构三句话:一服务一 crate、机制一文档一归属、两处同步即故障源。这个仓库用 22 个 crate 和 11 篇设计文档实现了一台消费级机器人该有的工程秩序,而它的"大文件"——8 千行 main 与 6 千行协议——恰好标记出系统的两根承重柱:控制循环与契约。

第 3 章 进程间通信:Unix socket 上的 JSON-RPC 契约

本章代码走读:duck-ipc-proto/src/lib.rs(约 6200 行)——Id、Call、Service 三个类型撑起整个通信体系。

3.1 一个 crate 防一种死法

duck-ipc-proto 的存在理由写在 Call 枚举的文档注释里:

/// A method together with its parameters.
///
/// Every request is built from one of these and read back as one, so a method can never be
/// paired with another method's parameters — the drift this crate exists to prevent.
#[derive(Debug, Clone, PartialEq)]
pub enum Call {
    /// Version handshake. The first call on a connection.
    Hello(HelloParams),
    ...

方法与参数在类型层面绑定:不存在"方法名字符串对了、参数结构旧了"的请求。这是"契约即代码"的 Rust 表达——协议演化时,编译器替你找出所有没跟上的人。

3.2 代码走读:Call 枚举——整个系统的 API 表

约 100 个变体按命名空间分组,注释里直接标注了每个调用的节奏类型(这对理解系统时序至关重要):

    // ── intents ──────────────────────────────────────────────────────────────
    /// Continuous. Send as a notification.
    RobotMove(MoveParams),
    /// Continuous. Send as a notification.
    RobotHead(HeadParams),
    /// Discrete. Send as a request; the answer is [`LookResult`].
    RobotLook(LookParams),
    RobotStop,
    RobotEnable(EnableParams),
    /// Power the joints and ramp to the home pose. No policy needed.
    RobotInit,
    /// Cut power to the joints. The robot collapses if nothing holds it.
    RobotRelax,
    /// Reboot servos (all of them, or the ids named), then limp. See [`method::ROBOT_REBOOT_MOTORS`].
    RobotRebootMotors(RebootMotorsParams),
    /// Run a one-shot skill, or toggle sit↔stand.
    RobotDo(DoParams),
    /// Standing body pose. Continuous. Send as a notification.
    RobotPose(PoseParams),
    /// Mouth opening. Continuous. Send as a notification.
    RobotMouth(MouthParams),

连续调用是通知(notification),离散调用才是请求(request)——50 Hz 的遥操作指令走 fire-and-forget,丢一帧无所谓(下一帧 20 ms 后就到);配对、安装、配置这类低频高语义操作才需要应答。这个区分直接决定了 Id 的设计:

/// Request identifier. `None` on a [`Request`] makes it a notification.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(untagged)]
pub enum Id {
    Number(u64),
    Text(String),
}

命名空间的完整清单(每组的语义):

命名空间 职责 代表调用
update.* 组件更新与回滚 Check Apply Rollback ResetToGolden Select Pin Subscribe
robot.* 运动意图与生命周期 RobotMove RobotDo RobotInit RobotRelax RobotPolicies
policy.* 策略库管理 PolicyCheck PolicyInstall PolicyFetch PolicySearch
detector.* 视觉检测模型 DetectorCheck DetectorInstall
account.* 设备账号 AccountLogin AccountStatus AccountLogout
net.* Wi-Fi NetStatus NetScan NetConnect NetForget
system.* 系统信息与安全 SystemInfo SystemLogs SystemPairingPin SystemAuthenticate
pad.* 手柄 PadPair PadBindings PadInput
流订阅 传感器/媒体流 TofStream HeadImuStream

注意 TofStream 的注释:"Subscribe to the ToF depth stream. Answered by tofd."——谁应答写在契约里,路由不是猜的。

3.3 代码走读:Service 枚举——没有中间人的路由表

/// The service that owns the answer to a call.
///
/// One socket per service, connected directly — there is no broker (`architecture.md` §2.2). A
/// transport adapter holds connections to the services whose calls it carries, and to no others:
/// `btd` holds three, and `padd` being absent from them is deliberate rather than incidental —
/// `padd` is the unprivileged client whose whole value is having no special access.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum Service {
    /// `updaterd`, at [`DEFAULT_SOCKET`].
    ...

三个决策一目了然:

  1. 每服务一个 socket,直连,无 broker。没有 DBus、没有消息中间件——每少一个组件就少一类故障与一个依赖。
  2. 传输适配器只持有它需要的连接。最小权限不是口号:btd(蓝牙桥)只连三个服务。
  3. padd 的"无特权"是被设计出来的。手柄服务故意不持有任何特殊连接——它的全部价值就是"没有任何特殊访问"。攻击面管理进了类型系统。

3.4 代码走读:配对 PIN——一条无法按 BLE 规格实现的需求

配对认证是这个系统里最精巧的一段协议设计。SystemPairingPin:

    /// Read the pairing PIN.
    ///
    /// Exists so `btd` can answer a BLE passkey request without owning config. It must never be
    /// routed to BLE — a PIN an unpaired peer can read authorises nothing — and `btd`'s routing
    /// table has a test saying so.
    SystemPairingPin,

SystemAuthenticate:

    /// Prove knowledge of the robot's pairing PIN.
    ///
    /// Answered by the **transport** rather than by any service, which makes it unlike every
    /// other call here. BLE cannot express a fixed, printed-on-the-robot passkey — the spec has
    /// the *displaying* side generate a random one, and a headless robot can display nothing — so
    /// the PIN check moved from the link layer to this one, where we define the rules. See
    /// `docs/design/app-path-design.md` §5.
    SystemAuthenticate(AuthenticateParams),

拆开看:BLE 规格假定配对双方至少一方有屏幕(显示侧生成随机码);一台无头机器人既不能显示也不能遵循该模型,于是团队把 PIN 校验从链路层上移到应用协议层,并配套两条纪律——PIN 读取调用永不被路由回 BLE(否则未配对端读走 PIN,认证形同虚设),且这条禁令由 btd 路由表上的一个测试看守。又是第 2 章那招:"两处必须同步/禁止的事,用机制不用记忆。"

3.5 鸭群合唱:通知的三方协作

chorale(合唱)功能让多只鸭子互相听见、同步鸣叫,它的协议是三个调用的小型状态机,注释把方向讲得极清楚:

    /// `btd` subscribing to what it should advertise. Answered, then a stream of
    /// [`method::CHORALE_BEACON`] notifications.
    ChoraleSubscribe,
    /// `robotd` telling `btd` what to advertise. A notification.
    ChoraleBeaconSet(ChoraleAdvertise),
    /// `btd` telling `robotd` what it heard. A notification.
    ChoraleHeard(ChoraleHeard),

robotd 决定广播什么 → btd 通过无线电发出去并听别的鸭子 → 听到的内容回流给 robotd。感知-决策-通信的闭环全部走同一套 JSON-RPC 通知机制,没有为"多机协作"新造任何基础设施。

3.6 为什么不是 gRPC / DDS

v0.1 讨论过候选方案的表格,现在可以补上实证:Call 枚举里连续调用是通知、Service 直连无 broker、PIN 认证在传输层应答——这三个需求在 gRPC/DDS 里都要绕着框架做。而 6200 行的 duck-ipc-proto 用纯类型就把"方法-参数绑定、服务归属、节奏类型、路由禁区"全部编码完毕,serde 序列化后即为线上格式,jq/nc 即可调试。当你的 IPC 语义已经复杂到需要"契约"时,把契约做成类型比做成配置更有利。

3.7 本章小结

duck-ipc-proto 用三个枚举(Id、Call、Service)回答了 IPC 的全部四个问题:谁能被调用(方法表)、带什么参数(类型绑定)、谁应答(服务归属)、以什么节奏(通知 vs 请求)。第 4 章进入这些调用的最大消费者——robotd 的 50 Hz 控制循环。

第 4 章 robotd:50 Hz 控制循环与技能调度

本章代码走读:robotd/src/control.rs(约 750 行,纯计算核心)、robotd/src/main.rs(约 8100 行,IO 与常量)、duck-control/src/obs.rs(观测组装)。

4.1 一图流:一个 tick 里发生什么

control.rs 的模块注释是全仓库最好的一份"程序说明书",值得整段精读:

//! Turning sensors and a command into joint targets — and scheduling the skills.
//!
//! Everything here is pure computation between [`duck_control::io::RobotIo::read`] and the
//! safety layer's `apply`. It holds no IO handle — by construction it cannot command a
//! motor, only propose targets.
//!
//! The tick, in order:
//!
//! ```text
//! skill windows ← advance / expire (roulade window, kick timer, ground-pick phase, sit↔stand rise)
//! command      ← the caller's smoothed command, re-encoded for the active skill
//! net          ← roulade > kick > ground pick > sit/rise > stand-by-magnitude > walk
//! action       ← ONNX
//! targets      ← home pose + action_scale × action
//! filters      ← optional first-order low-pass on head and legs
//! ```

第一段划定了可测试性边界:控制器不持有任何 IO 句柄,"按构造它无法命令电机,只能提议目标"。读传感器 → 纯计算 → 安全层执行,三段切开后,中间那段可以在任何机器上做单元测试(第 6 章的仿真也受益于此)。

末行公式是策略输出的完整语义:关节目标 = 家姿态 + action_scale × 策略输出。策略学的不是绝对角度,而是围绕家姿态的偏移——这个约定必须训练与部署两侧一致。

4.2 代码走读:Tuning——每个数字都对着训练侧

pub struct Tuning {
    /// Scales raw policy output before it becomes a joint offset. The prototype's current
    /// alpha default.
    pub action_scale: f64,
    /// The standing policy is trained to be applied whole.
    pub standing_action_scale: f64,
    /// Standing runs softer, at this fraction of the running gain. `--standing-kp-ratio`.
    pub standing_gain_ratio: f64,
    pub gain: u16,
    /// First-order low-pass on the head joints. `None` is no filtering. The alpha policies
    /// are trained with 0.5 — it must match training or transfer degrades.
    pub head_lowpass: Option<f64>,
    /// Same, for the ten leg joints. Trained with 0.7.
    pub legs_lowpass: Option<f64>,
}

impl Default for Tuning {
    fn default() -> Self {
        Self {
            action_scale: 0.9,
            standing_action_scale: 1.0,
            standing_gain_ratio: 0.8,
            gain: 200,
            head_lowpass: Some(0.5),
            legs_lowpass: Some(0.7),
        }
    }
}

注意 head_lowpass: Some(0.5) 的注释——"alpha 策略是以 0.5 训练的,必须与训练一致,否则迁移退化"。低通滤波不是部署侧可以自由调节的"手感参数",而是训练契约的一部分(RL 仓库 AGENTS.md 从另一侧说了同一件事:策略默认无滤波,单方面在任一侧加/改滤波都会破坏迁移)。两个仓库在这类细节上的互文,就是 sim2real 的"契约面"。

4.3 代码走读:网络选择优先级链与相位编码

step() 的后半段是全章最密的逻辑——根据当前状态选择驱动网络,并把状态编码进命令槽:

        } else if let Some(phase) = self.ground_pick {
            // The twist slots carry the phase encoding; head and body are zero-padded,
            // mirroring the training env's `zero_command_padding`.
            let angle = std::f64::consts::TAU * phase;
            let c = Command {
                twist: [angle.cos(), angle.sin(), 0.0],
                ..Command::default()
            };
            (Net::GroundPick, c, "ground_pick".into())
        } else {
            let mut c = *command;
            match self.sit {
                // The posture flag rides the twist vx slot: 1 = sit, 0 = stand. Head and
                // body slots stay live — the prototype keeps them in the buffer too.
                Sit::Sitting => {
                    c.twist = [1.0, 0.0, 0.0];
                    (Net::SitStand, c, "sit".into())
                }

两处"盗用"命令槽的编码:

  1. ground-pick 的进度相位被编成 (cos 2πφ, sin 2πφ) 塞进 twist 的 vx/vy 槽——用两个正交分量表达一个循环相位,策略能从观测里无歧义地解码"拾取动作进行到哪一步",且 φ=0 与 φ=1 在圆上是同一点(天然周期连续)。
  2. 坐/站姿态标志骑在 twist vx 槽上:1 = 坐,0 = 站。注释特意说明头和身体槽保持活跃。

这类"一槽多用"是对 61 维统一观测的极致利用:不改布局,就能让不同任务族在同一个向量里携带各自需要的信号。

优先级链本身(模块注释里的 roulade > kick > ground pick > sit/rise > stand-by-magnitude > walk)在代码中体现为 if-else 的顺序;注释还保留了两个容易"好心修坏"的遗产:

//!  - **A kick window runs at standing tuning.** The kick's observation carries an all-zero
//!    command, and in the prototype the standing transition fires on exactly that — so a
//!    kick runs at `standing_action_scale` and the softened standing gain. Kept, because
//!    the kicks were tuned against it.

踢球窗口的观测命令全零,恰好触发了"切站立调参"的分支——踢球实际跑在站立的柔和增益下。这是缺陷还是特性?团队的选择:保留,因为踢球就是在这个条件下调出来的。对行为系统而言,"历史一致性"有时比"逻辑纯洁性"更值钱。

4.4 代码走读:COAST_TICKS——一次"随机微小抽搐"事故的尸检

main.rs 里最值得全文摘录的注释,讲的是如何对待例行故障:

/// **One dropped Dynamixel transaction is ordinary** — `robotd` says so itself when it logs
/// one — and it must therefore be *invisible*. It was not: a failed read used to stop the
/// policy for that tick, which commanded the hold pose and reset the controller, so every
// ordinary dropped read produced a visible twitch. Measured on a bench robot at ~8 drops a
/// minute with a monitor attached, which is exactly the reported "random tiny spasms".
///
/// Coasting is safe because the observation is *already* a tick old by construction: at
/// 50 Hz the policy is trained on data of exactly this age, and a second tick of it is
/// inside that. Three ticks (60 ms) covers a drop, a retry and a slow tick; past that the
/// robot genuinely cannot see, and holding still is the honest answer.
const COAST_TICKS: u32 = 3;

完整的问题-诊断-论证-参数选择链条:

配套的 RESET_AFTER_PAUSE(200 ms)同理:控制器复位本身是不连续动作,20 ms 的网络抖动不配触发它,200 ms 的真中断才配。故障处理的粒度要和故障的真实频率对齐——这是嵌入式控制系统的通用课。

4.5 代码走读:观测组装的两条禁令

第 1 章看过 61 维布局表,obs.rs 随后列出了"容易做错、且是对着原型逐项确认过而非想当然"的两条:

//!  1. **Body x, y and yaw are hardcoded zero.** They are unbound in the training
//!     environment, so an all-zero body command is the *nominal* encoding, not a
//!     placeholder standing in for something better.
//!  2. **Head targets ride in the command, and are not added on top of the policy output.**
//!     The prototype does both, in different modes, and gates the post-hoc addition behind
//!     `if !new_cmd_obs` with the note "head_offsets are a COMMAND fed via the obs vector
//!     instead — don't double-add it here". Doing both would bend the head twice.

第二条尤其典型:头部目标是喂给策略的命令,不能在策略输出之后再叠加一次——"两处都做会把头掰弯两次"。跨系统接口的参数,"谁消费"必须唯一。

4.6 循环与 IPC 的关系

main.rs 的开篇注释交代了 8100 行 daemon 与控制循环的边界:IPC 侧"publishes and never calls into the loop"——只发布、绝不回调进循环;一个卡死的循环靠自己报告不健康("a wedged loop reports itself unhealthy"),而不是被外部探测杀死。日志预算同样体现纪律:每 tick 都记日志的话"50 Hz 一天约 430 万行",所以常态静默、每 5 分钟(LOOP_SUMMARY_INTERVAL)汇总一次,孤立丢包 60 秒内不重复抱怨。

4.7 策略侧:槽位、递归与热切换

从 Call 枚举(第 3 章)已经见过 RobotPolicies / RobotLoadPolicy / RobotReloadPolicies 与 PolicySlot 结构——robotd 按槽位管理策略(行走、站立、坐站、拾取、技能各占一槽),加载即热切换,且支持"重读全部槽位"。main.rs 注释还提到:"API 2 adds explicit recurrent state; API 1 feed-forward models remain supported"——循环策略(带隐状态)已在协议里占位,前馈模型继续兼容(详见 docs/recurrent-policies.md)。

4.8 本章小结

robotd 的工程设计可以提炼成四条:纯计算核心(可测试性按构造保证)、调参即契约(每个数字对着训练侧)、例行故障必须不可见(滑行三拍)、接口参数消费唯一(不许掰弯两次)。8100 行里最珍贵的不是算法,是这些把"上机才暴露的问题"提前写进代码纪律的注释。下一章快速巡览其余守护进程。

第 5 章 周边服务:更新、媒体、感知流与"个性"

本章代码走读:Call 枚举的 update.* 与 detector.* 组、docs/design/ 的更新与重启设计文档、kinematics 的性能化改造。

5.1 updaterd:更新是一套动词,不是一次动作

第 3 章的 Call 枚举里,update.* 是最大的命名空间之一,动词清单本身就是更新系统的状态机:

    // ── update.* ─────────────────────────────────────────────────────────────
    Check(ComponentParams),
    Apply(ApplyParams),
    Rollback(ComponentParams),
    ResetToGolden(ComponentParams),
    Select(SelectParams),
    Pin(PinParams),
    Status,
    ListInstalled(ComponentParams),
    Log(LogParams),
    /// One run's full transcript. `update.log` says which runs exist; this says what one did.
    Show(ShowParams),
    /// Turns the connection into a stream of [`method::PROGRESS`] notifications.
    Subscribe,

Apply/Rollback/ResetToGolden(恢复出厂金色镜像)/Select(选通道)/Pin(钉住当前版本)——每个动词对应一个真实的运维场景,进度通过 Subscribe 转成通知流。配套的设计文档有三篇:updater-design.md、restart-order.md(服务重启顺序——多守护进程系统里"谁先起谁后死"是要写下来的)、boot-recovery-net.md(变砖后的网络恢复路径)。

更新与健康门禁的联动藏在 robotd 的开篇注释里——main.rs 第一行就自报家门:

//! The 50 Hz control loop, the `robot.*` socket, and the health the updater gates on

robotd 对外提供 RobotHealth 与 RobotSafeToRestart 两个调用(第 3 章见过),而 updaterd 以它们为门禁:机器人正在走路时不可安全重启,更新引擎必须问过控制循环才能动手。健康不是仪表盘上的装饰数字,是更新流程的前置条件——"检测到不健康的组件就回滚"因此有了权威数据源。

策略的更新与守护进程的更新是两条独立通道(第 1 章 Cargo.toml 的 [workspace.metadata.policies]:官方策略集在 HF Hub,"步态重训不应要求守护进程发版"),但共享同一套哲学:任何能远程到达机器人的东西,都要有版本、校验和回滚。

5.2 mediad:WebRTC 主路径与检测模型联动

第 2 章从 AGENTS.md 读过媒体策略:消费者走 WebRTC(端到端加密、同会话携带控制信道、有回传路径),media.stream(WebSocket 推帧)只是程序消费纯帧流的兜底。这里补一个 Call 枚举里的联动细节:

    // ── detector.* ───────────────────────────────────────────────────────────
    /// What detector is installed and what the Hub offers; see [`method::DETECTOR_CHECK`].
    DetectorCheck,
    /// Install a detector and restart `mediad` onto it.
    DetectorInstall(PolicyInstallParams),

视觉检测模型(跑在 RK3566 的 NPU 上,duck-detect crate)与 RL 策略走同一套安装参数结构(PolicyInstallParams)和同一套 Hub 分发通道——"模型即资产"不区分它是神经网络检测器还是步态策略。安装检测器会"把 mediad 重启到它上面",媒体管线与感知管线在同一个进程里交汇。

5.3 tofd 与传感器流:订阅制

深度与 IMU 不在 Call 的请求-应答里,而是专门的流订阅:

    /// Subscribe to the raw pad input stream. Answered by `padd`, not `configd`.
    PadInput,
    /// Subscribe to the ToF depth stream. Answered by `tofd`.
    TofStream,
    /// Subscribe to the head IMU (BMI088 on the HAT); see [`method::HEAD_IMU_STREAM`].
    HeadImuStream,

注释顺带披露了硬件细节:头上那枚 IMU 是 BMI088,焊在 HAT 板上。流式接口与第 3 章"连续调用是通知"的节奏分类一脉相承——高频数据只订阅,不请求。

5.4 声音与个性:theremin、chorale 与一只鸭子的灵魂

robot.* 组里有三个"不像正经机器人功能"的调用,它们是产品人格的载体,也是工程上最讨巧的部分:

    /// Play a voice-bank sound.
    RobotSound(SoundParams),
    /// Pick the ToF theremin up or put it down. Discrete; the answer is [`ThereminResult`].
    RobotTheremin(ThereminParams),

ToF 特雷门琴:深度传感器测手距,映射为音高——传感器流(TofStream)的消费端除了避障还可以是乐器。配合 sounds crate 的语音库、chorale 的多鸭合唱(第 3 章),以及 pet-detect(宠物检测)驱动的互动行为,一只鸭子的"性格"由这些小模块拼出来。研究平台会把它们当 demo,产品公司把它们当核心资产——这个仓库的认真程度(theremin 有专门的 ThereminState 结构和结果类型)说明团队清楚自己在做后者。

5.5 代码走读:kinematics——为 50 Hz 而生的编译式 FK

kinematics/src/lib.rs 的模块注释是一次教科书级的性能化改造记录:

//! MJCF-driven forward kinematics, compiled for the control loop.
//!
//! The prototype's `microduck_kinematics` crate proved the approach — parse the
//! training MJCF, walk the tree — but its query path was built for convenience:
//! joint angles travelled in a `HashMap<String, f64>`, and asking for one site
//! recomputed every body in the robot into a freshly allocated `Vec`. Fine at
//! bench cadence; wasteful inside a 50 Hz loop that asks for both feet every
//! tick.
//!
//! Here the model is *compiled* once at load: names are resolved to indices,
//! and each site gets its own flattened root→site chain of links. A query is
//! then a fold over that chain — no hashing, no allocation, no bodies the site
//! does not hang from. Angles are a plain `&[f64]` indexed by [`Model`]'s joint
//! order; resolve names to indices once with [`Model::joint_index`], not per
//! query.
//!
//! Correctness is pinned two ways: `tests/fk_against_mujoco.rs` compares every
//! site against MuJoCo's own `mj_kinematics` on 64 random poses, and
//! [`head`] pins the sign conventions the rest of the system steers by.

叙事结构值得仿写:原型验证了路线(解析训练同源的 MJCF)→ 指出查询路径的两个罪状(字符串哈希、每次查询全树重算加分配)→ 给出编译式方案(加载时名字解析成索引、每个 site 一条展平的链、查询即折叠)→ 正确性双保险(对 MuJoCo 自身结果比对 64 个随机姿态 + 符号约定单独钉死)。"与 MuJoCo 逐点对齐"正是第 9 章 sim2real 哲学在几何层的回声:仿真用的模型与部署用的运动学必须同源互证。

5.6 本章小结

周边服务群展示了三条产品级思路:更新是带门禁的状态机而非单次动作;所有模型(检测器与策略)共用一条带校验的分发通道;高频数据一律订阅流。加上 theremin 与合唱这些"人格模块",一台消费级机器人的软件版图就此完整。下一章看这些服务如何在没有硬件的情况下整体运转起来。

第 6 章 duck-sim:没有机器人也能开发

本章代码走读:scripts/duck-sim(shell 脚本头注释)、docs/design/simulation.md。

6.1 一条命令一只鸭

scripts/duck-sim 是个 shell 脚本,它的头注释就是完整的用户手册:

#!/bin/sh
# A duck in MuJoCo, driven by the real daemon, in one command.
#
#   scripts/duck-sim              # a window opens, the duck stands up, and it is yours
#   scripts/duck-sim drive        # walk it forward for a few seconds
#   scripts/duck-sim ctl health   # anything robotctl does
#   scripts/duck-sim down
#
# And the same duck, as a machine you log into:
#
#   scripts/duck-sim boot         # the daemons under real systemd, in a container
#   scripts/duck-sim shell        # you are on the duck
#   duck-a # robotctl health
#
# `boot` needs sudo, because `systemd-nspawn` does. What it buys over the plain form is the part
# that is not fakeable: the real unit files, with their real User=, groups, RuntimeDirectory= and
# hardening, under a real init — which is what makes `robotctl update apply` and the health gate
# behave here the way they behave on a robot.

两级真实度,按需升级:普通模式起 MuJoCo 窗口与真实 robotd;boot 模式把守护进程装进 systemd-nspawn 容器,跑真实的 systemd 单元文件——真实的 User=、用户组、RuntimeDirectory= 与加固选项。注释说得很直白:单元文件的权限与沙箱配置"不可伪造",而正是它们决定 robotctl update apply 和健康门禁在仿真里是否与真机同构。容器还给了鸭子一个主机名(duck-a),四只鸭子各自登录、互相发现——多机系统的调试体验也有了。

6.2 设计文档:一只"孪生鸭"的边界

docs/design/simulation.md 开篇给出定位与实测数据:

**Status:** built and in use. `robotd --sim`, `tofd --sim` and `mediad --sim-camera` are here, with
`microduck_rl`'s `duck-body` serving the other half; the containers (§8), the ether (§5) and the
per-duck cameras all run from `scripts/duck-sim`. Measured: the daemon holds `50.0 of 50.0 Hz · 0
missed` against a MuJoCo body, detects a seated boot from the simulator's own joint angles, and
`robotctl robot init` runs the sitstand policy until the duck is upright and stays there; four ducks
in containers sing a full chorale over the ether. ...

The goal is a duck you develop against exactly as you develop against a robot: the same binaries,
the same units, the same `robotctl`, the same `duckctl open` — with the body in MuJoCo instead of on
the desk. Not a mock, and not a test harness. A twin, with a written-down boundary.

关键事实:

6.3 代码走读:三个"没人猜得出来"的坑

脚本头注释的最后一段,把脚本存在的理由总结为堵住三个坑(每一条都是环境差异的化石):

# The three things this exists to stop anyone typing by hand, because none of them is guessable:
#
#   * policies live at /opt/robot/daemon/current on a robot and in this repo on a laptop, so every
#     one has to be named in a params file;
#   * `ort` dlopens libonnxruntime and a laptop has no system one — the RL repo's venv does;
#   * a unix socket path is capped at ~108 bytes, so the state directory has to be short.
  1. 策略路径双轨:机器人上在 /opt/robot/daemon/current,笔记本上在仓库内——于是仿真参数文件必须逐个点名策略;
  2. ONNX Runtime 的动态加载(第 1 章伏笔回收):robotd dlopen 系统 libonnxruntime,笔记本通常没有——但 RL 仓库的 venv 里有,脚本去那里找。开发环境复用训练环境,一个实际的"仓库间依赖";
  3. Unix socket 路径约 108 字节上限:每服务一 socket(第 3 章)的架构下,状态目录必须短——这是内核限制倒逼的目录命名设计。

把"不可猜的坑"封装进一条命令,是开发者体验(DX)工程的核心:不是文档写得细,是让人根本不需要读文档。

6.4 与 RL 训练仿真的分工

duck-sim(主仓库) mjlab 环境(RL 仓库)
目的 软件在环:验证守护进程、协议、部署链路 大规模并行训练(4096 环境)
运行的代码 真实 Rust 守护进程(--sim 开关) 纯 Python 训练环境
身体服务方 RL 仓库的 duck-body mjlab/MuJoCo Warp
频率 实时,实测 50.0/50.0 Hz 零丢拍 GPU 加速,非实时
逼真度侧重 系统行为(systemd、更新、门禁) 执行器/齿隙等物理保真

两个仿真不是竞争关系而是一条验证链的两段:RL 仓库训出策略 → duck-sim 验证"真守护进程吃得下" → 真机。RL 仓库的 scripts/ 里还有专门的 sim2real 对比脚本,把同一条策略在两侧的轨迹并排量差。

6.5 无硬件学习路线(修订版)

结合 Sandbox 与 duck-sim,零硬件的完整路径:

  1. 第 0 天:Hugging Face Sandbox 浏览器里建立手感;
  2. 第 1 周:scripts/duck-sim 起鸭,robotctl 遍历命令;再 boot 进容器看 systemd 单元与健康门禁的真实行为;
  3. 第 2–3 周:读 duck-ipc-proto 与 robotd/src/control.rs,在仿真里给自己的 RPC 方法或技能写代码——改 control.rs 是安全的,它无 IO、可单测;
  4. 第 4 周起:进入 RL 仓库(下一章),训练并用 duck-sim 验证自己的策略。

6.6 本章小结

duck-sim 的启示不在 MuJoCo,而在三件纪律:孪生要写明边界、真实度分级可选(普通/容器)、把环境差异的坑封装成一条命令。"同一套真软件,两种身体"是第 3 章契约设计的最大红利,也是这个仓库对"机器人开发难"给出的最系统性的回答。至此工程篇结束,下半部进入策略的产地。

第 7 章 强化学习基础与 microduck_rl 的工作流

本章代码走读:RL 仓库 AGENTS.md 的命令与仓库地图、tasks/ 目录、microduck_velocity_env_cfg.py 头部。

7.1 技术栈三层楼

microduck_rl 站在三层成熟基础设施上:

概念上只需带走 PPO 的三要素:观测(61 维,第 1/8 章)、动作(14 维关节位置偏移,joint_pos_action.scale = 1.0)、奖励(数十项加权和,第 8 章逐项走读)。训练与部署都以 50 Hz 节拍运行——频率是两侧共同的呼吸。

7.2 代码走读:AGENTS.md 命令表——工作流即纪律

RL 仓库的 AGENTS.md 命令区是团队的肌肉记忆清单,每行都有讲究:

## Commands

```bash
uv run list-envs                                    # live task registry
uv run train <TASK_ID> --env.scene.num-envs 4096    # train (add --hf-jobs for Hugging Face Jobs)
uv run train <TASK_ID> --env.scene.num-envs 64 --agent.max_iterations 5   # SMOKE TEST — always run first
uv run play <TASK_ID> --wandb-run-path <entity/project/run_id>
uv run scripts/export.py <TASK_ID> --wandb-run-path <...>   # → ONNX (bakes obs normalizer — mandatory path)
uv run publish --task <TASK_ID> --wandb-run-path <...> --checkpoint N --repo <user>/microduck-<name> --kind episodic --duration-s 4.0
                                                    # → HF Hub repo (policy.onnx + schema-2 manifest.json + README) the daemon loads via `robotctl policy add`
uv run scripts/infer_policy.py --walking out.onnx   # CPU MuJoCo deployment rehearsal (BAM M6 actuators as in training; --no-bam = XML PD)
uv run --with pytest pytest tests/

A 5-iteration smoke test at 64 envs catches ~95% of config errors for cents. Never launch a long run without one.

划线句:**5 迭代 × 64 环境的冒烟测试,花几分钱抓住约 95% 的配置错误;不冒烟不跑长训**。RL 实验最大的浪费不是 GPU 时费,是 90 分钟后才发现奖励项拼错字的一天。命令表还埋着完整链路的顺序:`train → play(回看)→ export(唯一合法转 ONNX 的路)→ infer_policy(CPU 部署彩排)→ publish(Hub)`,最后由机器人侧 `robotctl policy add` 消费——第 10 章逐环展开。

`infer_policy.py` 的注释点出一个细节:彩排用的执行器是 **BAM M6,与训练一致**(`--no-bam` 才退回 XML 自带 PD)——连"验证部署"这一步都在守护执行器保真度契约。

## 7.3 代码走读:仓库地图——一个 cfg 一个任务族

`AGENTS.md` 的 Repo map 节(节选):

```markdown
- `src/mjlab_microduck/tasks/mdp.py` — ALL custom MDP functions (rewards, events,
  observations, commands, curricula). Add new functions here, grouped by task.
- `src/mjlab_microduck/tasks/microduck_*_env_cfg.py` — one cfg module per task
  family. `microduck_velocity_env_cfg.py` is the main walking recipe AND the
  shared base (robot, DR, obs, commands) other envs build on or mirror.
- `src/mjlab_microduck/tasks/__init__.py` — task registration (base + `-Backlash-` variants).
- `src/mjlab_microduck/tasks/backlash.py` — wraps any env cfg into its backlash twin.
- `src/mjlab_microduck/robot/microduck_constants.py` — robot cfgs, HOME frame, BAM actuator cfg.
- `src/mjlab_microduck/robot/microduck/` — MJCF exports from Onshape
  ... Collision families: `walk` (feet only), `groundcontact` (curated floor set),
  `allcollisions` (every part; XL330 housings named `*_servo_collision` by
  `name_servo_collision_geoms` → VelStand's servo-impact sensor). Each has a
  `_backlash` twin generated by `add_backlash.py <xml> --backlash-deg 2.0`.

三件结构设计:

  1. mdp.py(约 7400 行)收拢全部自定义 MDP 函数——奖励/事件/观测/命令/课程一个文件按任务分组;新函数只进这里。大而集中换来的是"找一个奖励项的实现不需要猜它在哪个文件"。
  2. velocity 配置是主配方兼共享基座。行走任务的环境配置同时是其他任务族的模板:机器人配置、域随机化、观测、命令全部继承或镜像它。AGENTS.md 后文解释了为什么:"Building on make_microduck_velocity*_env_cfg keeps DR / obs / noise / delays in sync for free"(免费保持同步)——否则你要手工搬运整套 DR + 观测噪声 + NaN 防护栈。
  3. 碰撞族三分 + 齿隙孪生。walk(只算脚,最快)、groundcontact(精选接地部位)、allcollisions(全部部件,XL330 舵机外壳命名为 *_servo_collision,供 VelStand 的"舵机撞击传感器"用)——同一具身体三种碰撞精度,按任务需要选;每族再由脚本生成带 ±2° 齿隙的孪生版,-Backlash- 变体因此能和基础任务做无混杂的 A/B 对照。

任务族清单(tasks/ 目录 14 个 env cfg):velocity(行走)、velstand(行走+恢复)、standup(起身)、sitstand(坐站)、ground_pick(俯拾)、ball_kick(踢球)、roulade(前滚翻)、rollers/swizzle/spin/crouch/slope/roller_standup(轮滑六变体)、testbench(台架)、distill(蒸馏)。

7.4 代码走读:velocity 配置的"开关面板"

每个环境配置文件顶部是一排大写常量开关——这份列表本身就是域随机化的菜单,注释里还留着调参史:

# Domain randomization toggles
ENABLE_COM_RANDOMIZATION = True
ENABLE_HEAD_COM_RANDOMIZATION = True  # Randomize CoM of the head assembly bodies
ENABLE_KP_RANDOMIZATION = False #  Was True
ENABLE_KD_RANDOMIZATION = False #  Was True
ENABLE_MASS_INERTIA_RANDOMIZATION = True  # Can enable once walking is stable
ENABLE_JOINT_FRICTION_RANDOMIZATION = True  # Scales BAM's friction budget per-env via FrictionDRBamActuator.friction_scale
ENABLE_JOINT_DAMPING_RANDOMIZATION = False
ENABLE_ARMATURE_RANDOMIZATION = True  # Reflected rotor inertia (microban-style). DOES affect BAM (armature is set, not zeroed).
ENABLE_VELOCITY_PUSHES = True  # Velocity-based pushes for robustness training
ENABLE_IMU_ORIENTATION_RANDOMIZATION = True  # Simulates mounting errors
ENABLE_ENCODER_BIAS = True  # Per-env joint encoder calibration offset (actor obs sees joint_pos + bias)
ENABLE_BASE_ORIENTATION_RANDOMIZATION = False  # Randomize initial tilt to force reactive behavior

读点有三:ENABLE_KP_RANDOMIZATION = False # Was True——被关掉的开关留着"曾是 True"的注释,这是负结果的记录(增益随机化试过、撤了);ENABLE_MASS_INERTIA_RANDOMIZATION 的注释写着"行走稳定后可开"——随机化强度跟着策略成熟度走;ENABLE_ENCODER_BIAS 让 actor 观测看到"关节角 + 标定偏差"——仿真里就把真实编码器的装配误差灌进观测。这些开关的具体实现散在 mdp.py 与 friction_dr_bam.py,第 9 章逐个走读。

7.5 无 GPU 怎么办

训练要 CUDA(4096 环境约 1–2 小时出可用步态),但整条学习路径并非必须有卡:--hf-jobs 把训练提交到 Hugging Face Jobs;infer_policy.py 与 tests/(cfg 不变量与 MDP 函数回归测试)全部 CPU 可跑。冒烟测试本身也只要几分钱级别的算力。

7.6 本章小结

microduck_rl 的工作流三律:先冒烟再长训、新任务从最近模板长出来、随机化强度跟着策略成熟度走。基础设施(mjlab/rsl_rl)都是社区成熟的轮子,这个仓库真正的产出是"配方"——下一章进入配方的核心:奖励与观测。

第 8 章 microduck_rl 解剖:观测布局与奖励配方

本章代码走读:tasks/symmetry.py(观测布局权威文档)、tasks/microduck_velocity_env_cfg.py(主配方)、tasks/mdp.py(函数库巡礼)。

8.1 代码走读:symmetry.py——61 维的"第二份真理"

第 1 章读过部署侧(obs.rs)的 61 维布局表;训练侧的权威版本写在 symmetry.py 的模块注释里——这个文件的本职是实现左右对称数据增强,注释顺手把布局和关节序固化成了文档:

Actor observation layout (61-dim flat tensor, concatenated in term insertion order):
    [0:3]   base_ang_vel      (roll, pitch, yaw  — body-frame IMU)
    [3:6]   projected_gravity (gx, gy, gz         — body-frame)
    [6:20]  joint_pos_rel     (14 joints, relative to default pose)
    [20:34] joint_vel_rel     (14 joints)
    [34:48] last_action       (14 joints)
    [48:51] twist command     (lin_vel_x, lin_vel_y, ang_vel_z)
    [51:55] head command      (neck_pitch, head_pitch, head_yaw, head_roll deltas)
    [55:61] body command      (x, y, z, roll, pitch, yaw deltas)

Joint ordering within each 14-dim block (from robot_walk.xml body tree):
    0: left_hip_yaw    5: neck_pitch    9:  right_hip_yaw
    1: left_hip_roll   6: head_pitch    10: right_hip_roll
    2: left_hip_pitch  7: head_yaw      11: right_hip_pitch
    3: left_knee       8: head_roll     12: right_knee
    4: left_ankle                       13: right_ankle

关节序是左腿 0–4、头 5–8、右腿 9–13 的交错布局——RL 仓库的不变量清单专门警告:在 roller/backlash 模型上被动关节还会再交错,"永远不要在 mdp 函数里硬编码关节索引",要用 _servo_joint_ids 这类帮手函数。

对称增强本身是精巧的数据工程:左右镜像时,腿交换后 hip_yaw/hip_roll 取负(偏航/侧滚轴在镜像下反向),hip_pitch/knee/ankle 也要取负——不是几何原因,而是家姿态本身左右用了相反符号约定(left_hip_pitch = +0.6, right_hip_pitch = -0.6)。连"哪些维度取负"这样看似几何的问题,答案都取决于训练约定的细节。

8.2 代码走读:不变量——"零填充,永不删槽"

61 维布局能支撑策略热切换(第 4 章),靠的是一条铁律,AGENTS.md 原文:

- **Obs layout is 61D (actor) and shared across the whole policy family** so
  policies are hot-swappable in the runtime: 48 base proprioception +
  13D command block `[twist(3), head_pose(4), body_pose(6)]`, in that order.
  An env that doesn't use a command slot ZERO-PADS it (keep the obs term,
  sample tiny ranges) — never delete a slot.

不用的命令槽零填充,绝不删除。这样行走策略(用 twist)、坐站策略(用姿态标志)、拾取策略(用相位编码)共享同一个输入接口,运行时才能即插即换。第 4 章看到的"ground-pick 相位编进 twist 槽、坐站标志骑在 vx 槽"正是这条不变量在部署侧的镜像——61 维是全部策略族的公共插座。

8.3 代码走读:奖励配方——每一项权重都是一次实验

velocity 环境的奖励组装段是全书信息密度最高的代码。逐项走读:

姿态项排除头颈关节(正则的教训):

    # Pose reward operates on LEG joints only. Head/neck are command-driven
    # (head_pose_tracking) — if they were in this reward too, it would pull
    # them to HOME while head_pose_tracking pulls them to the command, and the
    # policy converges to "ignore the command" because pose reward dominates
    # once head_pose_tracking's gradient dies at large commands.
    cfg.rewards["pose"].params["asset_cfg"] = SceneEntityCfg(
        "robot", joint_names=(r"^(?!passive_|.*neck.*|.*head.*).*",)
    )
    cfg.rewards["pose"].weight = 1.0

两个奖励项争夺同一组关节(一个拉向家姿态、一个拉向命令),策略的理性解是"听大头、忽略命令"——奖励冲突不是叠加,是抵消。解法用一条负向前瞻正则把头颈从 pose 项里剔除。

直立项的定量论证:

    # upright: deliberately strong (2.0 / std²=0.05, was 1.0 / std²=0.1).
    # 2026-07 pitch-vs-speed eval: the policy walks with a +2-4° steady forward
    # lean (p90 ~6-8°) and ~2/3 of push-induced falls at speed are FORWARD. At
    # weight 1.0 / std²=0.1 a 4° lean cost ~0.05/step — effectively free. At
    # 2.0 / std²=0.05 it costs ~0.19/step: enough gradient to hold the trunk
    # level in steady gait while transient lean (push recovery, accel) stays
    # affordable.
    cfg.rewards["upright"].weight = 2.0
    cfg.rewards["upright"].params["std"] = math.sqrt(0.05)

一段完整的"测量 → 定价 → 调价":评估发现策略带着 +2–4° 前倾走路、2/3 的推撞跌倒是向前跌;旧参数下前倾 4° 每步只罚 0.05——"等于免费";加倍收紧后罚 0.19/步,"足够让躯干在稳态步态里保持水平,而瞬态前倾(推撞恢复、加速)仍可负担"。奖励权重不是超参数,是价格体系,调权前先算清每个行为现价多少。

故意留弱的足底打滑项:

    # foot_slip deliberately weak (-0.1, not -1.0): -1.0 was too restrictive
    # for this robot's pivot-heavy turning.
    cfg.rewards["foot_slip"].weight = -0.1

这只鸭子转弯靠 pivot(原地碾转),足底打滑是转弯方式的一部分,罚重了转不了弯。

原地转的采样修补:

# Fraction of envs commanded to spin on the spot (lin=0, |ang| ∈ [0.4·max, max]).
TURN_IN_PLACE_FRACTION = 0.15

头注释交代了病因:"independent uniform sampling makes spin-on-the-spot ~2% of data → untrained"(2026-07 审计)——速度与角速度独立均匀采样时,"线速度为零、角速度大"的原地转只占数据 2%,策略压根没学过。修法不是改奖励,是改指令分布:15% 的环境强制分派原地转命令。RL 调参的一半功夫在检查"策略到底见过什么"。

8.4 代码走读:头部下垂之战——一次教科书级的奖励手术

2026 年 8 月 20 日的修复记录,是整个仓库最精彩的一段注释,完整还原了一次失败的直接修法与成功的迂回修法:

    # Head droop fix (2026-08-20). The head walks pitched ~15° down (measured:
    # run ww1g2198 head_pose_tracking 1.544/2.0 → 14.6° mean joint error).
    # DO NOT fix this by tightening head_pose_tracking's std: run 5yay13u4 tried
    # fine_std=0.1 and the policy stopped walking entirely by iter 300 (air_time
    # 1.01 → 0.02, peak foot height 15 mm → 2 mm, entropy collapsed 10.9 → 1.9).
    # An instantaneous tight tolerance taxes walking 0.77/step — 76% of the whole
    # air_time reward — and is UNESCAPABLE, since a 280 g head (38% of robot
    # mass) must oscillate while stepping. Standing still scored higher, so it
    # stood still.
    # The DC bias, unlike the oscillation, IS escapable (bias the neck command up
    # to cancel gravity sag), so price only that: L1 on a 1 s EMA of the error.
    # At the optimum this costs a walking policy nothing.
    cfg.rewards["head_pose_bias"] = RewardTermCfg(
        func=microduck_mdp.head_pose_bias_penalty,
        weight=0.0,  # ramped by the head_pose_bias_weight curriculum below
        params={"command_name": "_head_pose", "tau_s": 1.0},
    )

逐句拆解这个案例:

  1. 症状量化:走路时头前俯约 15°;run ww1g2198 的头部追踪得分 1.544/2.0,折合平均关节误差 14.6°。注意他们用 wandb run id + 指标数字记录问题——可复现是调试的前提。
  2. 直接修法及其死亡:收紧 head_pose_tracking 的容差(std=0.1),run 5yay13u4——策略 300 迭代后彻底不走了:空中时间 1.01 → 0.02,峰值抬脚高度 15mm → 2mm,策略熵 10.9 → 1.9(坍缩)。
  3. 死因分析:瞬时紧容差对行走课税 0.77/步——占全部空中时间奖励的 76%——且这笔税"不可逃避",因为 280 g 的头占整机质量 38%,迈步时必然振荡,瞬时误差永远存在。站着不动得分反而更高,于是策略学会站着不动。"DO NOT fix this by tightening std"——开头这句禁令就是给未来的人省下 retrain 一次的学费。
  4. 迂回修法:区分可逃避与不可逃避的误差——振荡不可逃避(不罚),直流偏差可逃避(抬头颈命令抵消重力下垂即可消除),于是只对误差的 1 秒指数滑动平均(EMA,τ=1s)收 L1 罚。最优点上,走得好且头不垂的策略付税为零。

这一段的教学价值超过十篇论文:奖励设计的对象不是"错误",是"可逃避的错误"。对不可逃避的误差收税,等于教策略放弃整个任务。

8.5 mdp.py 函数库巡礼

7400 行的 mdp.py 收拢约 80 个函数,函数名即奖励配方全景(节选分组):

两个系统级细节:_nan_safe_reward_compute 的存在说明 NaN 不是异常而是要常态化防御的工况(GPU 并行仿真里一个环境发散不能烧掉整批梯度);分部位、分状态的惩罚变体说明"平滑"从来不是一个项,是一族项。

8.6 本章小结

奖励配方的三条心法,全部来自上面的代码:权重是价格,调权先算现价(upright);奖励冲突会抵消,同组关节只能有一个主人(pose 正则);只罚可逃避的错误(head_pose_bias 的 EMA 手术)。观测侧的心法只有一条但最硬:61 维是全策略族的公共插座,零填充、永不删槽。下一章把这些配方放进 sim2real 的战场。

第 9 章 Sim-to-Real:执行器、齿隙与随机化的纪律

本章代码走读:actuator/friction_dr_bam.py(全文 112 行,两代执行器类)、robot/microduck/add_backlash.py、AGENTS.md 不变量清单。

RL 仓库开篇一句话立纲:"Sim2real transfer is the whole point"——本章的三段代码就是这句话的全部工程实现。

9.1 代码走读:FrictionDRBamActuator——给 BAM 开摩擦随机化的口子

"""BAM actuator with per-env friction-magnitude domain randomization.

The canonical ``bam.mjlab.BamActuator`` exposes per-env gain scaling (kp/kd) but
no friction hook, and under BAM MuJoCo's ``dof_frictionloss`` is zeroed in
``edit_spec`` (BAM computes friction itself in ``compute()``). So the stock
``dr.dof_frictionloss`` is a no-op here.

This thin subclass adds a per-env ``friction_scale`` that multiplies BAM's
velocity-INDEPENDENT friction budget (Coulomb + Stribeck + load-dependent) inside
``_compute_friction_budget`` — the term that carries the dominant sim2real
friction uncertainty (stiction / gearbox). The viscous (velocity-proportional)
term is left at nominal; scale it too by overriding ``compute`` if ever needed.

Non-accumulating: ``friction_scale`` is reset to 1.0 then set to a fresh sample
each episode by the ``randomize_bam_friction`` event (see tasks/mdp.py).
"""

三层信息:

  1. 静默 no-op 的陷阱:BAM(Black-box Actuator Model)在初始化时把 MuJoCo 的 dof_frictionloss 清零、改由自己算摩擦——所以拿现成的 dr.dof_frictionloss 域随机化去"随机化摩擦",什么都不发生,也不报错。这是域随机化最阴的故障形态:你以为在练鲁棒性,其实策略从没见过摩擦变化。AGENTS.md 把它列为不变量:"joint-friction DR must scale the actuator's friction_scale — dof_frictionloss is zeroed under BAM, so randomizing it is a silent no-op."
  2. 只随机主导项:摩擦预算里"速度无关"的部分(库仑 + Stribeck + 负载相关)承载了主要的 sim2real 不确定性(静摩擦/齿轮箱),按环境随机化;粘性项(速度比例)留名义值。随机化的刀口要落在不确定最大的量上。
  3. 非累积:friction_scale 每集先复位 1.0 再采新样——第 9.4 节的"累积事故"正是这条纪律的由来。

9.2 代码走读:BacklashEncoderBamActuator——编码器在齿隙的哪一侧

第二个类解决一个微妙得多的问题:舵机固件的 PD 环读到的角度,在齿隙的哪一侧?

class BacklashEncoderBamActuator(FrictionDRBamActuator):
    """FrictionDRBamActuator whose firmware PD reads the encoder THROUGH backlash.

    Backlash models (robot_groundcontact_backlash.xml) put an unactuated
    ``passive_<joint>_backlash`` hinge in series with each servo joint: the
    servo joint is the motor output, the backlash joint is the play between it
    and the link, and the link angle is their sum.

    On the real servo the magnetic encoder sits on the OUTPUT side of that
    play, so the firmware position loop closes on main+backlash — while the
    servo winds through the dead zone the measured position (and hence the PD
    error) doesn't change. This subclass reproduces that: ``cmd.pos`` fed to
    BAM's voltage control law becomes qpos[main] + qpos[backlash].

    ``cmd.vel`` is left motor-side on purpose: in BAM it drives back-EMF and
    friction, which are rotor physics, not an encoder-derived firmware signal.

    Degrades to a plain FrictionDRBamActuator on models without backlash
    joints (per-joint mask), so it is safe to use on any microduck model.
    """

代码本体只有一处改动,却精确复刻了真实舵机的控制结构:

    def get_command(self, data) -> ActuatorCmd:
        cmd = super().get_command(data)
        pos = cmd.pos + data.joint_pos[:, self._backlash_joint_ids] * self._backlash_mask
        return dataclasses.replace(cmd, pos=pos)

要点拆解:真实 XL330 的磁编码器在齿隙的输出侧,固件位置环闭合于"主关节 + 齿隙"之和——舵机在死区里转时,被测位置(进而 PD 误差)纹丝不动。仿真里把 cmd.pos 换成 qpos[main] + qpos[backlash],就再现了"电机转了、编码器没动"的死区行为。而 cmd.vel 故意留在电机侧:它驱动反电动势与摩擦计算,那是转子物理,不是编码器信号。同一行代码里两个量走不同侧——保真度建模要细到"每个信号在物理上来自哪里"。最后,对无齿隙模型按关节掩码优雅降级,一个类通吃全部模型。

9.3 代码走读:add_backlash.py——从 CAD 到 XML 的齿隙注入

齿隙在 MJCF 里的实现是每个舵机关节后串一个无驱动被动铰链:

"""Inject gearbox-backlash joints into an onshape-to-robot MJCF export.

For every actuated servo joint (``class="chosen_actuator"``) this inserts an
unactuated hinge on the same body / same axis right after it:

    <joint axis="0 0 1" name="left_hip_yaw" ... class="chosen_actuator"/>
    <joint axis="0 0 1" name="passive_left_hip_yaw_backlash" class="backlash"/>

The composite link rotation is main + backlash: the main joint is the servo
output (BAM drives it), the backlash joint is the play between the servo and
the link, free to wander within ±(backlash/2).

Naming: the ``passive_`` prefix means the new joints are automatically excluded
by every existing regex in the task configs (actuators ``^(?!passive_).*``,
joint obs, pose reward).

``--backlash-deg`` is the TOTAL peak-to-peak play (what you measure wiggling
the horn with the servo held); the joint range is symmetric ±deg/2.
"""

三处巧思:命名即排除——passive_ 前缀让所有现存正则(执行器、观测、姿态奖励的选择器)自动跳过新关节,注入脚本零改动兼容全仓库;参数以测量定义——--backlash-deg 2.0 是"握住舵机晃动输出喇叭测到的总峰峰值",关节范围取对称 ±1°,参数的物理含义直接对应一种手持测量操作;接触参数按仿真步长定制——默认 solref (0.02,1) 在这种小范围关节上会让限位超程约两倍,改用 0.01 = 2×sim_dt(dt=0.005 时"最 stiff 的稳定设置"),solimp 提高阻抗让齿面接触近乎刚性。

为什么对 ±1° 这么较真? 因为齿隙是"结构已知的误差",用它建被动铰链 + 编码器侧反馈,策略在训练时就体验过死区里"指令动了、反馈没动"的因果结构。对比第 8 章的 head droop:随机化对付的是"不知模型的残差",结构化建模对付的是"知道结构的误差"——能用物理结构表达的,不要用噪声表达。

9.4 代码走读:不变量清单——用事故写成的军规

AGENTS.md 的 "Invariants — do not break these" 一节,每条背后都是一次翻车:

- **Obs normalization is ON** → the normalizer must be baked into the ONNX.
  `scripts/export.py` does this; in-sim play hides the bug (it applies the
  normalizer anyway), so never hand-convert a checkpoint.
- **Policies are UNFILTERED** (no action low-pass in training). Don't add EMA
  filtering without a matched runtime flag and a transfer test — trained-with /
  deployed-without (either direction) breaks transfer.
- **Domain randomization must not accumulate across resets.** mjlab 1.3.0's
  `dr.*` ops with `operation="add"/"scale"` are natively non-accumulating (they
  re-read compile-time defaults); custom DR functions must restore-then-apply.
  An accumulating CoM randomizer once degraded every long run for months.
- If an obs is remapped to a sensor view (backlash encoder, bias), any tracking
  REWARD on the same quantity must measure the same view — otherwise the policy
  is punished for correcting what it sees.

四条军规的通用教训:

  1. 会自动修复 bug 的工具最危险:仿真回放(play)自带归一化,于是"手转 checkpoint 忘了归一化器"的 bug 在仿真里隐形、上机才爆——所以唯一的导出路径必须强制(第 10 章)。
  2. 滤波是双侧契约:训练加/部署不加(或反之)都破坏迁移——第 4 章 Tuning 注释的另一半。
  3. 随机化必须可复位:累积式 CoM 随机化"曾让每次长训劣化数月"——域随机化的 bug 不在单步,在跨 episode 的漂移,regression 测试要跨 reset 检查。
  4. 观测与奖励必须同视图:观测重映射到齿隙编码器视角后,追踪奖励也得量同一视角——否则"策略因纠正自己看见的东西而被罚"。

9.5 工作流守则:训练之前先验证物理

AGENTS.md 的 "Building a new env" 工作流里,最省钱的一条排在第二位:

2. **Verify physics assumptions in sim BEFORE training** — this is the single
   biggest time-saver:
   - A target/rest pose must be a stable equilibrium: hold its ctrl for 3 s
     from noisy inits and check TILT, not just height (a settle test that only
     records z reports fallen states as "resting fine").
   - Measure target heights off the actual robot in sim (e.g. trunk z under a
     standing policy), never carry them across model revisions. A 5 mm-wrong
     STAND_Z once turned the goal into an impossible target for days.

两个具体的坑:settle 测试只记录高度 z,会把"已经摔倒躺着"误报为"静置良好"——躺着的鸭子高度也挺稳定,必须查倾角;STAND_Z 差 5 毫米曾让目标变成物理上不可能到达的位置,"白耗数天"。训练前的物理验证是单一最大的时间节省——奖励调不通时,先怀疑目标本身不可达。

9.6 验证阶梯(修订版)

结合两个仓库的真实工具链,从训练到上机的阶梯是:

GPU 训练(mjlab,BAM 执行器)
  → play 回放(wandb run 复查指标:air_time、熵、追踪得分)
  → CPU infer_policy.py(独立运行时彩排,BAM M6 与训练一致)
  → duck-sim 软件在环(真守护进程 + MuJoCo 身体,验证部署链路)
  → 真机(保守指令起步)

每一级排除一类问题:回放排除"训练代码自身 bug",CPU 彩排排除"训练框架依赖",duck-sim 排除"守护进程与协议 bug",真机首跑只承担物理世界的残余差距。RL 仓库 scripts/ 里的 sim2real 对比脚本则专门量化"同一条策略两侧轨迹差多少"。

9.7 本章小结

Microduck 的 sim2real 三板斧,每一斧都有代码为证:执行器当被控对象建模(BAM + 摩擦随机化开在正确的口子上 + 编码器位置精确到齿隙哪一侧)、结构化误差显式建模(被动铰链 + 命名即排除 + 按步长定制接触参数)、契约用机制守护(唯一导出路径、非累积随机化、观测-奖励同视图)。外加一条元纪律:训练之前先验证物理。下一章走完最后一环:ONNX 导出与 Hub 分发。

第 10 章 从训练到上机:ONNX 导出、manifest 与 Hub 分发

本章代码走读:export.py 头注释、publish/manifest.py(形状门禁与来源记录)、Call 枚举 policy.* 组、Cargo.toml 的 [workspace.metadata.policies]。

10.1 代码走读:export.py——唯一的合法路径

导出模块的开篇注释把"为什么只能走这条路"讲透了:

"""Export a trained checkpoint to ONNX, with the observation normalizer baked in.

This is the ONE path from a checkpoint to a deployable `.onnx`: `runner.export_policy_to_onnx`
emits `actor(normalizer(obs))`, so what the robot runs is what training saw. In-sim `play`
applies the normalizer itself and hides a hand-converted checkpoint that forgot it — never
convert by hand.

`scripts/export.py` is the command-line wrapper; `mjlab_microduck.publish` calls
:func:`run_export` directly so a published policy cannot skip this step.
"""

三个设计互锁:

  1. 归一化器烘焙进图:导出产物是 actor(normalizer(obs)) 的复合图——机器人端只需要标准 ONNX 运行时,不存在"图对了、归一化没带上"的事故面;
  2. 禁手转的理由是具体的:不是"怕你转错",而是 play 工具会自动补归一化、把忘带归一化器的手转模型藏得好好的——上机才暴露。工具自动修 bug 的地方,就是 bug 能溜过测试的地方;
  3. 发布器直接调用导出函数:publish 不接受现成 ONNX 绕过导出——"发布的策略不可能跳过这一步"。路径唯一性靠调用图保证,不靠文档劝告。

10.2 代码走读:manifest.py——上传之前先替 daemon 拒一遍

发布侧的形状门禁,常量与校验逻辑:

#  ... 61 = 48 proprioception + 13 command; 14 = the servos.
OBS_LEN = 61
ACTION_LEN = 14

def check_onnx(path: Path) -> OnnxShape:
    """Refuse a file the daemon would refuse at load: wrong widths, or one that is not 61 -> 14."""
    ...
    if shape.obs_len != OBS_LEN:
        raise ManifestError(
            f"{path.name}: observation width is {shape.obs_len}, the robot builds {OBS_LEN} "
        )
    if shape.action_len != ACTION_LEN:
        raise ManifestError(f"{path.name}: {shape.action_len} actions, the robot has {ACTION_LEN}")
    return shape

注释一语道破架构:"拒绝一份 daemon 也会拒绝的文件"——上传侧预先执行设备侧的加载门禁。这样坏策略根本到不了 Hub,而不是装到一半在鸭子身上报错。文档注释还注明契约文档的位置:"Contract = docs/policy-manifest.md in the microduck repo"——两侧共享一份规范文档。

manifest(schema-2)还带 Provenance 区块——git_provenance 记录导出时的仓库状态、checkpoint 来源、wandb run 路径(ExportResult 的字段),让每个发布物可追溯到训练现场。发布命令同时限定可发布的策略类别:

uv run publish --task <TASK_ID> --wandb-run-path <...> --checkpoint N \
  --repo <user>/microduck-<name> --kind episodic --duration-s 4.0

AGENTS.md 的限定:"only constant-command episodic/perpetual policies are publishable (phase/posture-flag are the set's)"——社区可发布的是"常量命令"的策略(限时型 episodic 或持续型 perpetual),而相位编码、姿态标志这些命令槽盗用技巧属于官方策略集专用。这是把第 4/8 章那些"高级编码"与社区生态隔开的护栏:公共分发通道只走最不容易被用错的形态。

10.3 设备侧:槽位、安装与热切换

机器人侧的消费接口全在第 3 章读过的 Call 枚举里:

    // ── policy.* ─────────────────────────────────────────────────────────────
    /// What is installed and what the Hub offers; see [`method::POLICY_CHECK`].
    PolicyCheck,
    /// Install a set and make it live; see [`method::POLICY_INSTALL`].
    PolicyInstall(PolicyInstallParams),
    /// Fetch one policy into the library; see [`method::POLICY_FETCH`].
    PolicyFetch(PolicyFetchParams),
    /// Search the Hub; see [`method::POLICY_SEARCH`].
    PolicySearch(PolicySearchParams),

以及 robot.* 组的槽位管理:RobotPolicies(各槽位跑什么,返回 PoliciesResult,含每个槽的路径与来源)、RobotLoadPolicy(加载或复位某槽)、RobotReloadPolicies(全部重读)。PolicySlot 结构描述单个槽位——这就是第 4 章"热切换"的协议层入口。注意 PolicyInstall 的注释"Install a set and make it live"——安装即生效,没有"装完还要记得重启"的暗坑。

出厂策略集的版本政治在第 1 章的 [workspace.metadata.policies] 里:官方集在 pollen-robotics/microduck-policies,version = "v5" 是最低版本——新刷机的板子装到它,低于它的板子被安装后钩子拉上来,高于它的不动。步态升级走 robotctl policy update,不需要 daemon 发版。

10.4 端到端复现清单(v0.2 修订)

零硬件完整闭环,每一步都是真实命令:

# 0. 冒烟(纪律,不是可选)
uv run train <TASK_ID> --env.scene.num-envs 64 --agent.max_iterations 5

# 1. 正式训练(无本地 GPU 加 --hf-jobs)
uv run train Mjlab-Velocity-Flat-MicroDuck --env.scene.num-envs 4096

# 2. 回放复查(看 air_time / 熵 / 追踪得分,记下 run id)
uv run play <TASK_ID> --wandb-run-path <entity/project/run_id>

# 3. 导出(唯一合法路径,归一化器进图)
uv run scripts/export.py <TASK_ID> --wandb-run-path <...>

# 4. CPU 部署彩排(BAM M6 执行器与训练一致)
uv run scripts/infer_policy.py --walking output.onnx

# 5. 软件在环(真守护进程 + MuJoCo 身体)
cd ../microduck && scripts/duck-sim ctl "policy fetch <user>/microduck-walk"

# 6. 发布(上传侧先跑形状门禁与来源记录)
uv run publish --task <TASK_ID> --wandb-run-path <...> \
  --repo <user>/microduck-<name> --kind episodic --duration-s 4.0

# 7. 真机(有鸭子的话)
robotctl policy add <user>/microduck-<name>

10.5 给自研项目的分发模板

把这条链抽象成五条可移植的规则:导出自包含(归一化器/预处理进图,部署端零隐式状态);路径唯一(发布器直接调导出器,绕过在调用图上不可能);上传即门禁(设备侧会拒绝的,上传侧先拒绝);manifest 带来源(git/checkpoint/run 可追溯);生态分级(社区走安全形态,高级技巧留给官方集)。

10.6 本章小结

导出-发布-安装链用约三百行 Python 加几个枚举变体,实现了消费级 OTA 的全部关键性质:唯一、自包含、预校验、可追溯、可回滚(配合第 5 章的更新引擎)。全书两个仓库的故事至此闭环:Rust 世界负责"让策略安全地跑起来",Python 世界负责"让跑起来的策略值得安全地跑"。

第 11 章 动手与贡献:学习路线与带走的东西

本章代码走读:两个仓库的 AGENTS.md/CONTRIBUTING.md 中与上手直接相关的部分。

11.1 三种读者的路线图(v0.2,全部真实命令)

路线 A · 系统工程向(Rust)

第 1–2 周   Rust 基础;通读主仓库 README、docs/design/architecture.md、robotd-design.md
第 3 周     scripts/duck-sim 起鸭;robotctl 遍历命令;对照第 3 章精读 duck-ipc-proto 的 Call 枚举
第 4–5 周   精读 robotd/src/control.rs(750 行,纯计算可单测)与 duck-control/src/obs.rs;
            用 scripts/duck-sim boot 进容器,观察 systemd 单元与健康门禁
第 6 周起   挑 docs/project/ 里的已知问题或 CONTRIBUTING.md 的入门议题,提第一个 PR

路线 B · 强化学习向(物理 AI)

第 1 周   概念补课(PPO、域随机化);跑 uv run list-envs 认识任务注册表
第 2 周   冒烟测试 → 4096 环境训 Velocity(无卡用 --hf-jobs)→ play 复查
第 3–4 周 改配方重训:动一个奖励权重(先按第 8 章"算现价")或一个 DR 开关,
          对比 wandb 指标(air_time、熵、追踪得分)
第 5 周起 挑一个模板造新任务:episodic 特技→standup 模板;两态→sitstand;
          动作→roulade(读它的 cfg 文档注释——一条五次实验的教训链)
第 6 周起 走完 export → infer_policy → publish,发一条自己的策略到 HF Hub

路线 C · 产品/硬件向:Sandbox 体验 → 研读第 5 章的更新/配网/媒体设计(updater-design.md、app-path-design.md、remote-access-design.md)→ 对照自己的产品找差距。

11.2 贡献须知(来自仓库的原话)

主仓库 AGENTS.md 的规矩第 2 章已述(机制一文档一归属;行为与文档不一致时文档是 bug)。RL 仓库侧另有两条硬要求:

高价值贡献方向:新任务与奖励设计(附仿真与真机对照视频)、tofd 深度接入策略观测(社区最易做的增量实验——当前全部策略纯本体感知)、duck-sim 边界扩展、文档与 faq 改进。

11.3 带走五件事(v0.2,每件都有代码为证)

  1. 契约即类型,两处同步即故障源——Call 枚举绑定方法与参数、default-members 收拢二进制清单、ONNX Runtime 版本三方互锁(第 2、3 章);
  2. 例行故障必须不可见——丢一次总线包滑行三拍,因为观测本来就滞后一拍(第 4 章 COAST_TICKS);
  3. 权重是价格,只罚可逃避的错误——upright 先算每步现价再调倍数;head droop 对瞬时振荡免税、只罚 1 秒 EMA 的直流偏差(第 8 章);
  4. 结构化误差显式建模,随机化留给残差——齿隙用被动铰链加编码器侧反馈建模,摩擦随机化开在 BAM 正确的口子上(第 9 章);
  5. 分发链唯一、自包含、预校验——归一化器进图、发布器直调导出器、上传侧替 daemon 先拒坏文件(第 10 章)。

11.4 为什么这只鸭子不该接 ROS 2

读到这里的 ROS 2 用户多半会问同一个问题:这套 Rust 守护进程 + Unix socket 的活,用 ROS 2 不是现成的吗? 这个问题值得认真回答一次,因为答案不是"不能",而是三条算得清楚的账。

一个反常的对照

先看事实:Pollen Robotics 的科研平台 Reachy 2 原生跑 ROS 2 Humble,暴露 ros2_control、TF、robot_state_publisher——这是 Pollen 工程师 Remi Fabre 2025-04-24 在 ROS Discourse 的官方发布帖里写明的,LeRobot 的 reachy2 文档里那行 -e ROS_DOMAIN_ID 与挂载 ~/.ros/log 的 Docker 参数也是佐证。

同一家公司、同一套 Rust 底层、同一个团队的两种选择。 所以 Microduck 不用 ROS 2 不是能力问题,是定位问题:Reachy 2 是给研究者接 RViz、写论文、复用生态的;Microduck 是给消费者拆箱拿手柄就走的。判据不在技术,在"用户拿到它要干什么"。

四条算不过来的账

其一,硬件预算。 RK3566 + 1 GB RAM + 32 GB eMMC,同时要喂 ONNX Runtime、WebRTC、NPU 视觉检测。整套 Rust 工程约 30 个依赖(媒体报道口径)。ROS 2 不是跑不动,是它带进来的一整套发行版依赖、每节点数十 MB 常驻、DDS discovery 的组播与内存开销,在这块板子上每一样都要从别处挤。

其二,回路要的是确定性,不是互操作。 第 4.4 节的 COAST_TICKS = 3 是量着"一次丢包 + 一次重试 + 一次慢拍"定出来的 60 ms 容差,前提是一个进程、一个循环、不被抢占。往里塞 executor 回调、DDS 序列化、跨进程 hops,加的不是平均延迟,是尾延迟——而尾延迟正是"随机微小抽搐"的成因。

其三,契约已经做成类型了,中间件反而表达不了。 第 3.6 节讨论过 gRPC/DDS:连续调用是 notification、离散才是 request,谁应答写进 Service 枚举,连"PIN 绝不路由回 BLE"这种禁区都有测试看守。这套语义在 DDS 里要么表达不了、要么得绕着框架做。附带的好处很朴素:NDJSON over Unix socket,nc 和 jq 就能调,1 GB 的板子上没装 Foxglove 也能排障。

其四,ROS 2 不替你还产品债。 OTA 签名与回滚、PIN 配对、WebRTC、设备账号——这些第 5 章的活必须做,跟用不用 ROS 2 无关,用了还得多维护一层。

更关键:它在解决一个已经被解决的问题

桥接真正的价值在跨机——Unix socket 是机内的,出了板子就不存在。但 Microduck 的跨机通道已经有了:第 2.4 节那条守则说得明白,消费级走 WebRTC(端到端加密、控制与视频同一会话、有回传路径),media.stream 只是备用;开发者侧还有 duckctl over BLE,连 SSH/WiFi 都不用配。再开一条 DDS 通道,是替一道已解出的题写第二种解法。

至于 ROS 2 生态里最诱人的四样东西,鸭子一样也吃不下:

ROS 2 能给 Microduck 用得着吗 为什么
RViz / TF 可视化 否 TF 树就是两条腿加一个头,可视化价值接近于零
rosbag 录制 基本否 它是 RL 策略不是模仿学习——策略靠仿真域随机化练出来,不吃真机演示数据
Nav2 导航 否 媒体报道其 ToF 为 8×8(64 像素)深度,够避障,建不了图
多机协同 否 chorale 已经用无线电 + 三条 JSON-RPC 通知做完了(第 3.5 节)

什么时候才真的该接

三条判据同时成立再动手:

  1. 硬件在手——能实测尾延迟,而不是纸上推演;
  2. 能说出要复用的具体包名——是 ros2_control 还是 Nav2 还是某个驱动,不是"生态"两个字;
  3. 上层计算不在板上,且现有 WebRTC 通道扛不住——比如要 50 Hz 双向流式控制外加多传感器硬时间同步。

第 3 条是唯一可能成真的:RK3566 跑不动 VLM,想做"VLM 看一眼 → 指挥鸭子"必然跨机。但那条路 gRPC/WebSocket 也走得通,DDS 不是唯一解,更不是默认解。

真要接,也只接外围:写一个桥接节点作为 duck-ipc-proto 的又一个客户端(身份与 duckctl 同级),往外出 /joint_states、/imu、/tf、相机与 ToF,往里把 /cmd_vel 翻译成 RobotMove 通知、把技能与策略切换做成 service。并且守住三条纪律:只有 robotd 能写电机(ROS 侧发的是意图,不是指令);连续量走 notification、离散才走 request,别给 50 Hz 指令加请求-等待;61 维观测的组装权只属于 duck-control,桥里不许再算一遍(第 1 章那句"最高风险代码")。核心回路一行不改。

比桥接更值得做的一件事

如果真想在鸭子身上做点官方没做的东西,不是桥接,是数据回流管线:

真机 rollout → 落 61 维观测 + 14 维动作 + 关节电流/电压 + 同步视频 → 转成数据集推 HF Hub → 反哺仿真的执行器参数辨识。

第 9 章那批参数——库仑摩擦、齿隙 ±1°、电压-力矩曲线——目前是仿真里编出来的。真机录下来的电流与位置-指令误差是它们唯一的校准来源,而这条闭环官方尚未提供。它不需要 ROS 2(一个 duck-ipc-proto 订阅者加一个 parquet 写入器就够),价值却直接落在 sim2real 这一本书的主线上。

一句话

ROS 2 解决的是多团队、多进程、可复用组件之间的互操作;Microduck 的问题是单板、单闭环、单供应商。前者要松耦合,后者要确定性。当系统里只有一个 50 Hz 闭环时,中间件不是资产,是抖动来源。

判断力体现在"知道什么时候不用",而不在"知道怎么用"。

本节来源:Reachy 2 的 ROS 2 栈——Remi Fabre(Pollen Robotics),ROS Discourse,2025-04-24;LeRobot 官方文档《Reachy 2》Docker 参数段。Microduck 硬件与依赖规模——Gear Live《Hugging Face's $399 Microduck》2026-08、cocoloop《399 美元开源机器鸭》2026-09。书内引用:COAST_TICKS(第 4.4 节)、chorale(第 3.5 节)、WebRTC 主备关系(第 2.4 节)、61 维观测(第 1.4 节)、执行器随机化(第 9 章)。

11.5 结语

v0.1 说"把本书当地图,把仓库当领土"。v0.2 写完后可以更进一步:这本书引用的每一段代码都来自两个仓库的真实文件,但仓库在流动——策略集已到 v5,61 维布局也曾是 51/54/49/85 维的演化终点。读的时候请保持第 2 章那条警惕:注释的嗓门不等于现状的优先级,行为的归属以设计文档为准。愿你训出的第一条步态,在第一次上机时稳稳落地。

附录 A:关键链接

附录 B:本书引用的主要源码文件

文件 行数 出现章节
microduck/duck-ipc-proto/src/lib.rs ~6200 1、3、5、10
microduck/robotd/src/main.rs ~8100 4
microduck/robotd/src/control.rs ~750 4
microduck/duck-control/src/obs.rs — 1、4
microduck/kinematics/src/lib.rs — 5
microduck/scripts/duck-sim — 6
microduck_rl/src/mjlab_microduck/tasks/symmetry.py 169 8
microduck_rl/.../tasks/microduck_velocity_env_cfg.py 949 7、8
microduck_rl/.../tasks/mdp.py ~7400 8
microduck_rl/.../actuator/friction_dr_bam.py 112 9
microduck_rl/.../robot/microduck/add_backlash.py — 9
microduck_rl/src/mjlab_microduck/export.py — 10
microduck_rl/src/mjlab_microduck/publish/manifest.py — 10
两仓库 AGENTS.md、主仓库 Cargo.toml、docs/design/*.md — 全书