第 2 章 总体架构:Rust workspace 与守护进程群

本章代码走读:Cargo.toml 的 default-members、AGENTS.md(主仓库)、docs/ 目录。

2.1 22 个 crate 的分工

主仓库是一个单 workspace,成员一眼可以分成四类:

类别 crate 说明
守护进程 robotd updater configd btd padd mediad tof 机器人上常驻,一服务一进程
领域库 duck-control duck-ipc-proto duck-detect duck-ether kinematics odometry pad-imu pet-detect sounds robotd-params uyvy 纯逻辑或纯协议,被守护进程与工具复用
客户端工具 robotctl duckctl 机上 CLI / 开发者笔记本 CLI
构建支撑 test-support xtask deploy(脚本) 测试夹具、自定义 cargo 子命令、部署

体量分布很说明问题:最大的单文件是 robotd/src/main.rs(约 8100 行,第 4 章主角)和 duck-ipc-proto/src/lib.rs(约 6200 行,第 3 章主角)——系统的心脏是"控制循环"和"契约"这两块,其余 crate 都小得多。控制器的核心计算被抽到 robotd/src/control.rs(约 750 行)保持可读,这是"大 daemon、小核心"的典型布局。

注意 tof(深度传感器)——博客里说的 tofd 守护进程对应的 crate 名是 tof;同样,更新服务 crate 名是 updater(daemon 二进制名 updaterd)。命名上"d"后缀标记 daemon 二进制,crate 名不带。

2.2 代码走读:default-members 与一台不该看到蓝牙栈的开发机

第 1 章读过 members,紧随其后的 default-members 有一段仓库里最生动的注释——duckctl(笔记本上的客户端)被排除在板机构建之外的原因:

# Everything except `duckctl`, and that exception is the whole reason this key exists.
#
# `cargo board --bins` — in `dev-push.sh` and in the release workflow — builds every default
# member for aarch64, which is right for a robot's daemons and wrong for a client a developer
# runs on their own machine: it would cross-compile a Bluetooth stack for a board that must never
# see one, on the release path, for nothing.
#
# The alternative was naming the binaries explicitly at each `--bins` call site. There are two of
# them and they would have to be kept in step by hand, which is the failure this repo keeps
# writing down. One list here, and a new daemon is picked up by both without anybody remembering.

一层意思浮在表面(省一次无用的交叉编译),另一层是方法论:"两处手工保持同步"被这个仓库视为必须用机制消灭的故障模式——要么收进一个列表,要么写个测试盯着。这个思想在后文反复出现:ONNX Runtime 版本三方互锁、PIN 路由表测试、策略 manifest 双侧校验,全是同一招。

2.3 文档体系:机制唯一归属

docs/ 的组织在这个仓库被上升到制度。主仓库 AGENTS.md 开篇立规矩:

## Docs own mechanisms; one page each

`docs/README.md` assigns every mechanism to one design doc. When a fact belongs to a page listed
there, every other page says one sentence and links. When two pages disagree, the one that does
not own the mechanism is the bug — and when behaviour and a design doc disagree, the doc is
the bug.

设计文档目录 docs/design/ 现有 11 篇,篇名即机制清单:

app-path-design.md     architecture.md        boot-recovery-net.md
policy-channel-design.md  remote-access-design.md  remote-webrtc.md
restart-order.md       robotd-design.md       simulation.md
updater-design.md      webrtc-console.md

另有 policy-manifest.md(策略清单契约——第 10 章的主角,且 RL 仓库的发布代码显式引用它)、recurrent-policies.md(循环策略 API)、faq.md(面向"对着一只鸭子开发"的人的任务型入口)。

对写书人的启示:这个仓库里注释与文档不是代码的附庸,而是决策记录。当行为与文档不一致时,被判死刑的是文档——因为腐化的文档比没有文档更危险。

2.4 三条产品级守则

主仓库 AGENTS.md 还有三条容易被研究型团队忽视的守则,值得全文摘录:

其一,消费级视频链路的主备关系:

## A consumer uses WebRTC. `media.stream` is the fallback

The robot publishes H.264 over WebRTC, and that is the default for anything consuming a duck's
camera — it is encrypted end to end, it carries the control channel on the same session, and it
has a return path. ...

This is worth stating because the repository reads the other way round if you only follow the
code: `media.stream` was built when the relay endpoint was dead and WebRTC genuinely could not
connect from a data centre, so its module doc argues its own case at length. That endpoint is
fixed. Do not conclude from the volume of prose that it is the preferred path.

——"不要因为某个模块的注释嗓门大就以为它是主路径"。代码库会留下历史的沉积层,读库要区分"现在时"与"过去时"。

其二,永不围绕版本差异做设计:老客户端的局限是拿来提版本号的理由,不是绕行的理由;版本偏差被记录并继续服务,只有真正缺失的路由或未知参数才可以拒绝。

其三,release 才是修复到达机器人的方式:"main being fixed is not a robot being fixed."——机器人走 stable 通道,修复要切 release 才到用户手里。研究代码没有这个概念,产品代码必须有。

2.5 守护进程模型回顾

对照第 1 章的表:robotd(控制平面)、updaterd(维护平面)、configd/btd/padd(连接与输入平面)、mediad/tofd(媒体与感知平面)。故障隔离、权限分离、独立演进的好处上一版已述;代价——IPC——是下一章整章的主题。

2.6 本章小结

架构三句话:一服务一 crate、机制一文档一归属、两处同步即故障源。这个仓库用 22 个 crate 和 11 篇设计文档实现了一台消费级机器人该有的工程秩序,而它的"大文件"——8 千行 main 与 6 千行协议——恰好标记出系统的两根承重柱:控制循环与契约。