KEEL · 龙骨 · A CURRICULUM FOR THE AI ERA
06. 一个动作五分钟后才完成怎么办? — keel 龙骨
发布、批量导入、视频处理和设备校准无法在一次短 HTTP 请求中完成。外部系统通常先返回“已受理”,再通过轮询、回调或事件报告最终状态。
发布、批量导入、视频处理和设备校准无法在一次短 HTTP 请求中完成。外部系统通常先返回“已受理”,再通过轮询、回调或事件报告最终状态。
受理不等于成功
{
"status": "pending",
"operation_id": "op-4821",
"accepted_at": "2026-08-22T10:00:00Z"
}
此时 Harness 应把 Command 标为 pending,保存 operation_id,释放当前请求线程。把 202 Accepted 写成 succeeded 会让用户看到虚假的完成状态。
Operation 状态机
stateDiagram-v2
[*] --> accepted
accepted --> running
running --> succeeded
running --> failed
running --> cancelling
cancelling --> cancelled
running --> unknown: 状态源不可达
accepted --> expired: 超过 deadline
“请求取消”与“已经取消”不同。外部系统可能无法中止已开始的动作,Harness 必须继续对账,防止迟到成功把已关闭 Run 悄悄改回成功。
轮询需要节制
轮询循环至少考虑:
- 使用外部系统建议的
Retry-After; - 指数退避并加入 jitter;
- 不超过 Command deadline;
- 每次查询都关联相同 operation ID;
- 状态不认识时进入
unknown,不要猜测; - Harness 重启后可以从持久化 operation ID 继续。
Webhook 或消息事件能降低轮询,但需要验证签名、防重放,并处理事件乱序和重复投递。
Worker lease 防止两个执行者同时接管
分布式 worker 领取任务时应获得有限期 lease:
lease_owner = worker-7
lease_until = 10:01:00
worker 定期续租;崩溃后 lease 过期,其他 worker 才能接管。旧 worker 恢复后提交的迟到结果必须带 lease/fencing token 检查,否则两个执行者可能同时推进同一任务。
运行实验
python courses/advanced/real-world-execution/course/project/examples/05_long_running_job.py
示例用确定性 Job Store 模拟 accepted -> running -> succeeded,不依赖 sleep。这样测试可以准确控制每次状态变化。
检查理解
- 外部系统返回 operation ID 时,本地必须保存什么?
- cancellation requested 为什么不是 cancelled?
- lease 过期后,为什么还要防止旧 worker 的迟到提交?