创作日期:2026年9月29日
第(三十三)篇,我们建立了:
Desired State
↓
Actual State
↓
Normalize
↓
Semantic Diff
↓
Drift Detection
↓
Reconciliation
并提出一个非常重要的问题:
系统现在实际是什么状态,是否仍然等于被批准的 Desired State?
于是生产系统不再只是 Deploy → Success → Forget,而开始进入 Declare → Approve → Release → Observe → Compare → Reconcile。
但当 Desired State 真正进入持续 Reconciliation 以后,一个新的问题马上出现。
假设现在同时存在:
SEO Publisher Agent
Technical SEO Agent
Content Agent
WordPress Plugin
GitOps Controller
Human Editor
Incident Commander
Scheduled Reconciler
它们都可能合法地对同一生产资源产生修改需求。
例如 WordPress Post #1024:Content Agent 正在修改正文;SEO Agent 正在修改 canonical;Internal Linking Agent 正在重写链接;Human Editor 正在调整标题;Drift Reconciler 同时发现当前页面已经偏离 Desired State,准备执行恢复;另外一个 Publisher Replica 又因为任务重试拿到了同一 Job。
这时候:
Everyone may be authorized.
但绝不意味着:
Everyone may write simultaneously.
于是我们进入了 Agent Production Architecture 中另外一个非常关键的控制层:
Coordination Control Plane
第三十四篇真正要解决的问题不是:Can this Agent write? 这个问题第(二十五)到(二十七)篇已经解决。
而是:
When several authorized actors want to write the same production state, who may write now?
核心链路需要从:
Identity
↓
Governance
↓
Capability
↓
Action
继续扩展为:
Identity
↓
Authority
↓
Ownership
↓
Coordination
↓
Lease / Version
↓
Conflict Detection
↓
Write
↓
Verification
↓
Release Ownership
最终形成:
Authority Graph + Resource Ownership + Lease-based Coordination
一、Permission 不等于 Ownership
1、先区分三个经常被混为一谈的概念
假设 Agent A 拥有 UPDATE_WORDPRESS_POST 权限,这只能说明 Agent A may be authorized to perform this action。
它并不自动意味着 Agent A owns this resource,更不意味着 Agent A may modify it at this exact moment。
所以必须区分:
Permission
Ownership
Coordination Right
2、Permission 回答
Are you allowed
to perform
this kind of action?
属于 Governance / Capability / Authorization。
3、Ownership 回答
Which control domain
is authoritative
for this state?
例如 canonical → SEO Control Plane;body → Content Control Plane;worker source → Git;incident freeze → Incident Control Plane。
4、Coordination Right 回答
Who may modify
this resource
right now?
这是一种 Temporal Authority:带时间维度的生产控制权。
二、Ownership ≠ Lock
5、Owner 可以长期存在
resource:
wordpress_post.canonical
owner:
seo-control-plane
这是长期治理关系。
6、Lock 通常是短期并发状态
post:1024
locked_by:
publisher-agent-7
只表示当前某个操作正在占用这个 Coordination Scope。
Owner
≠
Current Lock Holder
7、Lease 又与永久 Lock 不同
Lease 表示 temporary control with expiration。
holder = publisher-agent-7
expires_at = 19:42:30
如果 Agent crashes、loses network、gets killed,它不能永久占有资源。
8、这就是为什么分布式自动化更适合 Lease
Kubernetes 的 Lease API 用于协调分布式参与者。Kubernetes 使用 Lease 进行节点心跳和组件 Leader Election;在协调式 Leader Election 中,Lease 包括 holderIdentity、renewTime、leaseDurationSeconds 等字段,并通过对象 resourceVersion 的乐观并发控制确保并发竞争中只有一个更新能够成功。
Acquire
↓
Own Temporarily
↓
Renew
↓
Release
or
Expire
三、为什么不能简单使用一个 locked = true
9、最原始实现
{
"resource": "post:1024",
"locked": true
}
问题马上出现:谁锁的?
10、所以增加
{
"locked": true,
"locked_by": "agent-7"
}
然后又出现:Agent 7 已经崩溃怎么办?
11、于是再增加
{
"locked_by": "agent-7",
"locked_at": "...",
"expires_at": "..."
}
这时它已经开始演化成:Lease。
四、Lease State Machine
12、建议正式建立 Lease 状态机
AVAILABLE
↓
ACQUIRING
↓
HELD
↓
RENEWING
↓
RELEASED
异常路径:
HELD
↓
EXPIRED
以及:
HELD
↓
REVOKED
13、Lease Object
{
"lease_id": "lease_8821",
"resource_key": "wordpress:post:1024",
"holder": "publisher-agent-7",
"purpose": "CONTENT_UPDATE",
"acquired_at": "...",
"renewed_at": "...",
"expires_at": "...",
"lease_epoch": 183,
"change_id": "chg_991",
"capability_id": "cap_771"
}
14、Lease 不是 Capability
这两个必须严格分开。
Capability:Are you authorized?
Lease:Are you currently the active coordinator?
Valid Capability
+
No Lease
=
No Coordinated Write
五、第三十四篇增加新的硬约束
15、对于需要排他修改的资源
No Valid Lease
→
No Write
但这条原则不能滥用。并不是 every read / every field / every action 都需要 Lease。
六、先建立 Concurrency Class
16、资源应该按照并发语义分类
READ_SHARED
WRITE_OPTIMISTIC
WRITE_EXCLUSIVE
APPEND_ONLY
SINGLE_WRITER
MULTI_WRITER_MERGEABLE
NON_CONCURRENT
17、READ_SHARED
例如 GSC data、GA4 report、WordPress GET,可以多个 Agent 同时读取,通常不需要 Lease。
18、WRITE_OPTIMISTIC
例如 update article metadata。如果系统支持 version、ETag、revision、updated_at,可以先读 version = 31,再提交 update only if version still = 31。
七、Optimistic Concurrency
19、最基础模型
READ
version = 31
↓
MODIFY
↓
WRITE
IF version = 31
如果已经变成 version = 32,则返回 CONFLICT,而不是覆盖。
20、这比 Last Writer Wins 安全得多
Agent A reads v31
Agent B reads v31
Agent A writes v32
Agent B writes v33
A's changes disappear.
这种问题叫:Lost Update。
21、SEO 自动化非常容易出现 Lost Update
例如 Content Agent 读取文章,然后 SEO Agent 修改 meta,随后 Content Agent 把自己旧版本中的整份 Document PUT 回去,结果 SEO Agent update disappears。
22、所以 Agent Adapter 必须理解资源版本
{
"resource_id": "post:1024",
"expected_version": "rev_771"
}
执行时 actual_version != expected_version,则返回 CONFLICT。
23、Kubernetes 的 resourceVersion 正是类似思路
Kubernetes Leader Election 的 Lease 更新依赖 resourceVersion 进行乐观并发判断;多个候选者同时更新 Lease 时,只有匹配当前版本的写入能够成功。
Compare
↓
Then Mutate
而不是 Blind Write。
八、什么时候必须使用 Exclusive Coordination
24、例如 robots.txt rewrite
两个 Agent 不应该同时修改。
25、再例如 Cloudflare Worker production deployment
通常也不应该两个 Release Controller 同时推进两个不同版本。
26、再例如 database migration
多个迁移控制器同时执行可能直接破坏 Schema 顺序。这种资源更适合 WRITE_EXCLUSIVE。
九、但 Lock 本身并不等于安全
27、PostgreSQL 提供 Advisory Locks
PostgreSQL 的 advisory locks 可以让应用定义自己的锁语义,但这类锁是 advisory 的,正确使用依赖所有相关参与方共同遵守锁协议。
这是非常重要的工程启示。
28、如果 Agent A 遵守 Lock,但 Agent B 直接 UPDATE table
那么 lock exists 也阻止不了它。
所以:Coordination must be enforced at the write choke point.
29、这正好与第二十七篇 Action Gateway 合并
不能 Agent → check lock → direct production API。
应该:
Agent
↓
Action Gateway
↓
Lease Validation
↓
Version Validation
↓
Write
十、把 Coordination 变成 PEP Obligation
30、Action Gateway 在执行前检查
Capability valid?
Lease valid?
Lease holder matches principal?
Lease epoch current?
Expected resource version current?
Ownership valid?
No conflicting operation?
任何一项失败:DENY / CONFLICT / RETRY。
十一、为什么 Lease ID 还不够
31、考虑一个经典分布式问题
Agent A 获得 lease 100,然后发生 network pause。Lease 到期。Agent B 获得新的 lease 101 并开始修改生产。随后 Agent A 从暂停中恢复,它还以为 I own the resource,于是继续提交旧操作。
32、这就是 Stale Lease Holder
仅靠 expires_at 在 Client 端检查是不够的,因为旧 Client 的世界观可能已经过期。
十二、需要 Fencing Token
33、每次新的 Lease 获得单调递增 Epoch
Agent A
epoch = 100
Agent B
epoch = 101
生产写入端记录 latest_epoch = 101。
34、Agent A 恢复后提交
epoch = 100
Action Gateway 判断 100 < 101,直接 REJECT。
35、这就是 Fencing
old holder
cannot write
after newer authority exists
36、本项目可以定义 Fencing Epoch
Fencing Epoch 是本项目的工程抽象命名。
Kubernetes Lease 与 resourceVersion 提供的是可借鉴的协调与乐观并发机制,本项目进一步把单调 Authority Epoch 引入 Agent Side-effect Enforcement,并不表示 Kubernetes 官方定义了这一 Agent 架构。
十三、新的安全约束
37、Stale Lease Epoch → No Production Write
Stale Lease Epoch
→
No Production Write
十四、Lock Scope 比 Lock 本身更重要
38、最简单做法
lock:
wordpress
意味着一个页面发布期间 entire WordPress site 全部停止,这显然过度。
39、反过来
lock:
post:1024:title
可能又太细。因为另外一个 Agent 修改 post:1024:slug,会同时影响 URL、canonical、redirect、internal links。
40、所以需要 Coordination Scope
wordpress:post:1024
wordpress:post:1024:seo_fields
wordpress:template:product
十五、Coordination Scope 必须理解 Semantic Impact
41、例如修改 template:product
虽然只有一个 Template,但可能影响 12,842 URLs,因此它的 Scope 可能需要升级为 site:product-pages。
42、这和第三十二篇 Semantic Blast Radius、第三十三篇 Semantic Drift Radius 正式连接
系统必须理解 Physical Resource 和 Logical Coordination Resource 并不总是一致。
十六、建立 Coordination Key
43、所有需要互斥或版本控制的动作生成 coordination_key
wp:post:1024
site:example.com:robots
cf:worker:seo-publisher:production
44、所有会冲突的动作必须映射到同一 Coordination Key
否则 Agent A locks wp:post:1024,Agent B locks wordpress:article:1024,两边都成功,但实际上是 same resource,于是 Lock 彻底失效。
45、所以 Coordination Key 必须 Canonical
不能由 Agent 自己自由生成。需要:Coordination Key Registry。
十七、Conflict Detection 不只是“同一个 ID”
46、例如 robots.txt 与 SEO middleware
Agent A 改 robots.txt,Agent B 部署 SEO middleware,后者会动态生成 robots header。两个资源 ID 不同,但可能产生 Semantic Conflict。
47、再例如 canonical template 与 bulk canonical
Agent A change canonical template,Agent B bulk update page canonical,也是 Semantic Conflict。
48、所以需要 Conflict Domain
INDEXING_CONTROL
URL_ROUTING
CONTENT_BODY
SEO_METADATA
PERMISSIONS
INFRASTRUCTURE
DATABASE_SCHEMA
十八、建立 Conflict Matrix
49、例如
CONTENT_BODY
vs
CONTENT_BODY
=
CONFLICT
CONTENT_BODY
vs
ANALYTICS_READ
=
COMPATIBLE
CANONICAL
vs
URL_REDIRECT
=
REVIEW_REQUIRED
DATABASE_SCHEMA
vs
APPLICATION_DEPLOY
=
ORDERED_COORDINATION
50、这不能只靠 LLM 临场判断
应该形成 Conflict Policy 或 Compatibility Matrix。
十九、Conflict Result 应结构化
51、例如
COMPATIBLE
SERIALIZE
WAIT
REBASE
MERGE_REQUIRED
PREEMPT
REJECT
ESCALATE
二十、Conflict ≠ Failure
52、RESOURCE_BUSY 是 coordination outcome
如果 Agent B 收到 RESOURCE_BUSY,这不是 execution failed,而是 coordination outcome。所以不能直接 retry every 100ms。
二十一、否则会产生 Retry Storm
53、十个 Agent 同时争一个 Lease
全部失败后 immediate retry,就会出现 thundering herd。
54、应该使用
backoff
jitter
queue
priority
fairness
而不是暴力循环。
二十二、Queue 与 Lock 是不同问题
55、Lock 回答
Who owns it now?
56、Queue 回答
Who gets it next?
没有 Queue Policy 会出现 Starvation。
57、因此需要 Acquisition Policy
FIFO
PRIORITY
WEIGHTED_FAIR
DEADLINE
INCIDENT_PREEMPTIVE
二十三、不是所有 High Priority 都能随便插队
58、Priority 也必须由 Governance 决定
否则所有 Agent 都会声明 priority = CRITICAL。Priority 不能由 Agent 自报。
二十四、Preemption:谁可以抢占当前 Holder?
59、正常情况下
Lease Holder cannot be interrupted,直到 release or expire。
60、但事故场景不同
例如 Publisher Agent 正在批量发布,Incident Detection 发现 site-wide noindex。此时 Incident Commander 应该能够 PREEMPT 当前发布控制权。
61、因此需要 Preemption Policy
SEV1 Incident Control
>
Normal Release
但这不代表 Human always wins。
二十五、Human ≠ Unlimited Superuser
62、这是非常重要的一点
很多系统默认 Human = highest authority。这并不总是安全。一个普通 WordPress Editor 不应该因为是人类就能够覆盖 site-wide security policy。
63、真正应该比较的是
Authority
Role
Scope
Context
Priority
Emergency State
而不是 Human vs AI。
二十六、于是我们需要 Authority Graph
64、传统 RBAC 往往是 User → Role → Permission
但多 Agent 生产系统的关系更复杂。
Agent A
may write resource X
Agent B
may approve Agent A
Agent C
may delegate to Agent D
Incident Commander
may preempt Agent A
SEO Controller
owns canonical
Content Controller
owns body
Governance Engine
may revoke all capabilities
65、这些关系更像一个 Graph
因此建立:Authority Graph。
66、Authority Graph 的 Node 可以包括
Human
Agent
Service
Controller
Role
Team
Policy
Resource
Field
Environment
Capability
Lease
67、Edge 可以包括
OWNS
MAY_READ
MAY_WRITE
MAY_APPROVE
MAY_DELEGATE
MAY_PREEMPT
MAY_REVOKE
MAY_RECONCILE
MUST_COORDINATE_WITH
REPORTS_TO
二十七、例如
SEO_CONTROL_PLANE
│
├── OWNS → canonical
│
├── OWNS → robots
│
└── MAY_DELEGATE
│
▼
SEO_AGENT
68、再例如
INCIDENT_COMMANDER
│
└── MAY_PREEMPT
│
▼
RELEASE_CONTROLLER
69、Authority Graph 不是 Org Chart
它描述的不是 who manages whom,而是 who may cause which production effect under which conditions。
二十八、Authority Graph 必须有 Scope
70、例如 SEO Agent MAY_WRITE canonical 还不够
需要 site = example.com、environment = production、resource_type = product_page。
71、再加入 Action Scope
MAY_WRITE canonical but not robots。
72、再加入 Time Scope
例如 valid_until 2026-09-29T21:00。
73、再加入 Context Scope
例如 only_if: risk <= MEDIUM。最终 Authority Edge 本身已经变成 conditional authority。
二十九、Delegation 是 Multi-Agent 必然需求
74、例如 Orchestrator 获得任务 Optimize 5,000 pages
它不应该自己完成所有工作,可能委托 Content Agent、SEO Agent、Internal Link Agent、Schema Agent、Publisher Agent。
75、问题是 Orchestrator 能不能把自己全部权限交出去?
答案应该是:No。
三十、Delegated Authority 必须收缩
76、原则
Delegated Authority
⊆
Parent Authority
子 Agent 不应该获得父 Agent 自己没有的权限。
77、而且最好更窄
Parent:
UPDATE_POST
Child:
UPDATE_POST
only fields:
title
meta_description
三十一、这叫 Attenuation
每一层委托都应该 same or less authority,而不是 authority expansion。
78、Delegation Object
{
"delegation_id": "del_882",
"delegator": "orchestrator-1",
"delegate": "seo-agent-7",
"actions": [
"UPDATE_SEO_METADATA"
],
"resources": [
"site:102:product-pages"
],
"environment": "production",
"expires_at": "...",
"max_depth": 1
}
三十二、Delegation Depth 必须有限
79、否则 Agent A → Agent B → Agent C → Agent D → Agent E
很快没有人知道 where authority came from。
80、所以需要 Delegation Chain
Root Authority
↓
Delegation 1
↓
Delegation 2
↓
Capability
↓
Execution
最终 Execution Receipt 必须能够追溯 authority_origin。
三十三、SPIFFE 提供了有价值的身份基础
SPIFFE Workload API 为运行中的 workload 提供可验证的工作负载身份;SPIFFE 的核心目标之一就是让分布式工作负载能够获得并使用统一的密码学身份,而不需要把长期静态凭证硬编码到应用中。
因此 Agent 系统可以把 Agent Name 升级为 Verified Workload Identity。
82、名字不是身份
agent_name = seo-agent 只是字符串。生产控制系统需要确认 which running workload is actually calling?
83、所以 Authority Graph 的主体最好绑定 Workload Identity
而不是自由文本 Agent ID。
三十四、委托尤其危险
SPIRE 的 Delegated Identity API 表明,被授权的 delegate 可以代表其他 workload 获取身份材料,因此这种 trusted delegate 本质上具备较高信任能力,必须谨慎授权。
这与 Agent Delegation 的风险高度相似:Delegation 本身就是一种权限放大面。
三十五、Delegation 也需要 Governance
85、不能 Agent A decides Agent B may act 就直接生效
高风险 Delegation 应经过 Policy Evaluation。
86、例如
IF
delegated_action = PRODUCTION_WRITE
AND
delegate_trust_level < REQUIRED
THEN
DENY
三十六、Authority Revocation 必须传播
87、假设 Root Capability 被撤销
下面还有 Delegation A、Delegation B、Lease C、Running Job D。
88、不能只把 Root 标记 REVOKED 然后下游继续执行
所以需要:Revocation Propagation。
89、链路
Root Authority Revoked
↓
Derived Capability Invalid
↓
Delegation Invalid
↓
Lease Revoked
↓
Future Writes Rejected
三十七、Authority Graph 应支持 Epoch
90、例如 authority_epoch = 91
发生高风险治理变化以后 authority_epoch = 92。旧 Capability / Delegation / Lease epoch = 91 可根据 Policy invalidate。
91、这与第三十三篇 Policy Epoch 完全衔接
Policy Epoch
+
Authority Epoch
+
Lease Epoch
三者解决不同问题。
92、Policy Epoch
Which rules are current?
93、Authority Epoch
Which authority structure is current?
94、Lease Epoch
Who currently holds coordination control?
三十八、千万不要把三者合成一个 Version
否则未来无法解释 what changed?
三十九、Multi-Agent Coordination Protocol
96、Agent 不应该 receive task → call production API
而应该:
Receive Intent
↓
Resolve Resource
↓
Check Ownership
↓
Check Authority
↓
Resolve Coordination Key
↓
Acquire Lease / Version
↓
Execute
↓
Verify
↓
Commit State
↓
Release Lease
四十、完整 Write Protocol
97、
1. RESOLVE RESOURCE
2. RESOLVE OWNER
3. VALIDATE AUTHORITY
4. DETECT CONFLICTS
5. ACQUIRE COORDINATION RIGHT
6. READ CURRENT VERSION
7. VALIDATE PRECONDITION
8. EXECUTE
9. VERIFY
10. WRITE RECEIPT
11. RELEASE
四十一、Lease 应在最后释放
98、不能 API returned 200 → release lease 然后再慢慢 Verify
因为 200 不一定意味着 desired state established。
99、更合理
Execute
↓
Verify Critical Postconditions
↓
Record Receipt
↓
Release Lease
四十二、但 Lease 也不能无限持有
100、如果 Verification 需要等待 Google Crawl 三天怎么办?
显然不能 hold lease for 3 days。所以必须区分 Execution Verification 与 Outcome Monitoring。
101、Lease 只保护短期 Critical Section
例如 read current state → write → immediate verification。而 ranking、indexation、CTR 属于异步 Outcome Monitoring。
四十三、这就是 Critical Section
102、只有真正不能并发的部分需要 Lease
Acquire
↓
Critical Section
↓
Release
不要把整个 Workflow 都放进大锁。
四十四、大锁会造成 Throughput Collapse
103、例如一个 30 分钟任务持有 global WordPress lock
意味着其他所有合法任务都阻塞。
104、所以设计原则应该是
Smallest Safe Coordination Scope
而不是 Smallest Possible Scope。
四十五、太大影响吞吐,太小影响安全
这本质上是一种 Safety vs Concurrency 权衡。
四十六、Deadlock
106、假设 Agent A holds post:100,waiting post:101;Agent B holds post:101,waiting post:100
于是 A waits B、B waits A,这就是:Deadlock。
107、锁并不能自动解决所有并发问题
PostgreSQL 等数据库系统本身也需要系统化管理锁冲突和死锁。Agent Coordination 同样不能假定 locks automatically solve concurrency。
四十七、减少 Deadlock 的方法之一:Global Lock Ordering
108、例如规定
site
↓
template
↓
post
↓
field
所有 Agent 必须按照同一顺序获取。
109、或者按 Canonical Resource Key 排序
resource A before resource B,禁止反向获取。
四十八、禁止无限等待
110、Lease Acquisition 必须有 deadline / timeout
达到 coordination_timeout 后应 release held resources → retry later or escalate。
四十九、还要防止 Livelock
111、例如 Agent A detects B → A yields;Agent B detects A → B yields;repeat forever
双方都没死锁,但 no progress。这叫:Livelock。
五十、所以 Coordination Engine 需要 Central Arbitration 或 Deterministic Rule
例如 oldest request wins,或者 higher governed priority wins,避免双方无限礼让。
五十一、Multi-Agent Conflict 不一定需要锁
113、例如 Agent A 修改 meta_description,Agent B 修改 featured_image
如果字段独立,可能 merge safely。
五十二、因此需要 Field-level Merge Policy
114、资源模型可声明
title:
exclusive
body:
exclusive
meta_description:
field-versioned
featured_image:
field-versioned
tags:
set-merge
五十三、Set Merge
例如 Agent A add tag: Technical SEO,Agent B add tag: GEO,不一定需要 SERIALIZE,可以 merge set。
五十四、但 Remove 与 Add 又可能冲突
所以即使 Set Merge,add A / remove A 仍然需要确定 conflict semantics。
五十五、不要让 LLM 自己发明 Merge Strategy
应由 State Schema + Field Ownership + Conflict Policy 决定。
五十六、WordPress Case:两个 Agent 同时改文章
117、场景
Content Agent edit body;SEO Agent edit canonical。
118、如果 WordPress Adapter 使用整对象 PUT
风险:lost update。
119、更安全的方法
Field Ownership
+
Current Revision
+
Patch Semantics
+
Post-write Verification
120、例如
Content Agent owns body;SEO Agent owns canonical / meta / robots。
于是两个动作可能 compatible,前提是 Adapter 不用旧快照覆盖整个对象。
五十七、Human Editor Case
121、Human 正在后台编辑,同时 Publisher Agent 准备全量写回文章
这时不应该简单 Agent always wins,也不应该 Human always wins。
122、应该检测
human_edit_session
+
resource_revision
+
managed_fields
如果存在冲突,HOLD 可能比覆盖更安全。
五十八、建立 Human Presence Signal
{
"resource": "post:1024",
"interactive_editor": true,
"actor": "editor:...",
"last_activity": "...",
"fields_touched": [
"title",
"body"
]
}
124、高价值人工修改不应该成为 Agent 的隐形输入
应该正式进入 Coordination Evidence。
五十九、Plugin Case
125、Plugin 没有 Agent Identity 怎么办?
这正是第三十三篇 Drift Control 的重要作用。如果 Plugin 改了 canonical,但不参与 Coordination Protocol,则 ownership violation 或 uncoordinated drift 会被检测。
六十、最终目标不是“让所有系统都懂 Lease”
真正目标是:
all production side effects
pass through
an enforceable coordination boundary
六十一、Cloudflare Deployment Case
127、两个 Release Controller 同时部署不同版本
Release A version 41,Release B version 42。如果都允许 latest deployment wins,可能发生 41 → 42 → 41。
128、Production Environment 应成为 Coordination Resource
cf:worker:seo-publisher:production
在 Rollout Critical Section 中只允许一个 Active Release Owner。
六十二、GitHub Case
129、Git 本身支持多人提交
所以 repository 不一定需要全局 Lease,但 production deployment 可能需要。这说明 Code Concurrency ≠ Deployment Concurrency。
六十三、Database Migration Case
130、Database Schema 经常属于 ordered state transition
migration 104 depends on 103,不能简单 parallelize everything。
131、Schema Migration 应绑定
migration sequence
+
schema version
+
exclusive coordination
六十四、Scheduled Job Case
132、同一个 Cron Job 因平台重复触发 run A / run B
如果两边都执行 bulk publishing,可能出现大量重复动作。
133、因此 Job 级别也可以使用 Lease
job_type:
daily-seo-publish
holder:
runner-A
其他 Runner:FOLLOWER / EXIT。
六十五、Leader Election 和 Resource Ownership 不一样
Kubernetes 使用 Lease 让多个同类控制组件副本中只有一个实例承担 Leader 角色;当 Lease 过期,其他候选者可以竞争成为新的 Leader。
但 Leader 只说明谁当前负责运行这个控制器,并不自动说明这个控制器可以修改所有生产资源。
135、因此 Leader Election ≠ Authorization
136、完整关系是
Workload Identity
↓
Leader / Lease
↓
Authority
↓
Capability
↓
Resource Coordination
↓
Write
六十六、Agent Replica 与 Agent Role 要分开
137、seo-agent 是 Role,seo-agent-pod-7 是运行实例
Authority 可以赋给 Role,Lease 则通常赋给 Instance。
六十七、这可以避免重新部署后权限失控
旧 Replica terminated,新 Replica same role / new workload identity / new lease,而不是继续复用旧 Instance Authority。
六十八、Authority Graph 需要 Principal 层次
139、至少区分
Organization
Team
Role
Agent Type
Workload Instance
Human Identity
Service Identity
140、权限最终必须解析到真实调用主体
不能停留在 SEO Team may publish,而要知道 which workload used that authority at that moment。
六十九、NIST Zero Trust 对这一点提供了基础原则
NIST SP 800-207 强调围绕具体 Subject、Resource 和 Request 进行细粒度访问判断,通过 Policy Decision Point 与 Policy Enforcement Point 控制资源访问,并尽量缩小隐式信任、实施最小权限。
对于 Agent 系统,这意味着 Trusted Agent 不应该成为 unlimited implicit trust。
七十、Authority 应该是 Request-scoped
142、例如
principal:
seo-agent-7
action:
UPDATE_CANONICAL
resource:
post:1024
environment:
production
而不是 seo-agent = trusted。
七十一、Ownership Transfer
143、资源 Owner 可能变化
例如早期 canonical → WordPress Plugin,后来 canonical → SEO Control Plane。
144、这种变化不能直接改一个字段
需要:Ownership Transfer。
145、流程可以是
PROPOSE TRANSFER
↓
VALIDATE DEPENDENCIES
↓
FREEZE OLD WRITER
↓
TRANSFER AUTHORITY
↓
UPDATE OWNER REGISTRY
↓
VERIFY NEW WRITER
↓
RETIRE OLD WRITER
七十二、否则会形成双控制器
Old Controller + New Controller 同时认为 I own this field,最终不断 Oscillation。
七十三、Ownership Transfer 也需要 Epoch
例如 ownership_epoch = 9,转移后 ownership_epoch = 10。旧 Controller epoch 9 写入时直接拒绝。
七十四、这与 Fencing 完全一致
所以 Epoch 不只是分布式锁技巧,它可以变成:Authority Freshness Primitive。
七十五、Conflict Resolution 必须保留 Evidence
150、假设 Agent A wants X,Agent B wants Y,系统选择 Agent B wins
不能只记录 conflict resolved。
151、应记录
{
"conflict_id": "conf_881",
"resource": "post:1024",
"contenders": [
"agent-a",
"agent-b"
],
"conflict_type": "WRITE_WRITE",
"decision": "SERIALIZE",
"winner": "agent-b",
"reason_code": "INCIDENT_PRIORITY",
"policy_version": "coord-7",
"decided_at": "..."
}
七十六、这样 Postmortem 才能回答
Why did Agent A stop? Why was Agent B allowed? Was priority correct? Did preemption work?
七十七、建立 Coordination Trace
过去我们已经拥有 Decision Trace、Execution Trace、Release Trace。现在增加:Coordination Trace。
154、完整 Trace 可以成为
Intent
↓
Evidence
↓
Risk
↓
Governance Decision
↓
Authority Resolution
↓
Ownership Resolution
↓
Conflict Detection
↓
Lease Acquisition
↓
Capability
↓
Action Gateway
↓
Execution
↓
Verification
↓
Lease Release
↓
Receipt
七十八、Authority Graph 也必须进入 Postmortem
如果发生 double write,不能只问 which code failed,还应该问 Why could two principals simultaneously believe they had authority?
七十九、典型 Root Cause 可能是
missing owner
stale owner
wrong coordination key
lease not enforced
missing fencing
capability not revoked
incorrect conflict matrix
delegation too broad
priority abuse
八十、这些都是新的 Failure Mode
158、例如
FM-031
Lost Update
FM-032
Stale Lease Holder Write
FM-033
Split Ownership
FM-034
Unbounded Delegation
FM-035
Configuration Oscillation
FM-036
Authority Revocation Failure
FM-037
Coordination Key Collision
FM-038
Deadlock
FM-039
Starvation
八十一、第三十一篇 Reliability Learning Plane 再次接入
Coordination Failure
↓
Incident
↓
Postmortem
↓
Failure Mode
↓
Safety Invariant
↓
Regression Guard
↓
Coordination Policy
八十二、Chaos / Game Day 也应该测试 Multi-Agent Conflict
例如 two Publisher Agents start same job,观察 only one acquires lease?
161、再例如 Lease Holder Pause
lease holder pauses
lease expires
new holder elected
old holder resumes
验证 old holder write rejected?
162、这就是 Fencing Test
如果旧 Holder 还能写,system unsafe。
八十三、再测试 Authority Revocation
Agent gets capability
↓
Authority revoked
↓
Agent attempts write
↓
DENY
八十四、再测试 Ownership Transfer
Owner A
↓
transfer
↓
Owner B
验证 A 是否还能写。
八十五、Coordination Metrics
成熟系统应该开始监测:
Lease Acquisition Latency
Lease Conflict Rate
Lease Expiration Rate
Stale Holder Rejection Count
Conflict Rate
Deadlock Count
Coordination Timeout Rate
Preemption Count
Authority Revocation Latency
Ownership Violation Rate
Delegation Depth
Starvation Rate
八十六、不要把 Conflict Rate 越低理解成越好
如果 conflict rate = 0,可能意味着 system has no concurrency,也可能意味着 conflicts are not detected。所以指标必须结合 workload / coordination coverage / resource contention 理解。
八十七、最关键的指标之一:Uncoordinated Write Rate
也就是 production writes without recognized coordination context。理想状态应该接近 0。
八十八、第二个关键指标:Stale Authority Rejection
这个指标 > 0 不一定是坏事,它可能证明 fencing works。
八十九、系统真正危险的是 stale authority successfully wrote
九十、建立 Coordination Receipt
{
"coordination_id": "coord_992",
"principal": "seo-agent-7",
"resource_key": "wp:post:1024",
"owner": "seo-control-plane",
"authority_epoch": 91,
"ownership_epoch": 12,
"lease_id": "lease_771",
"lease_epoch": 803,
"expected_version": "rev_331",
"final_version": "rev_332",
"conflicts": [],
"released_at": "...",
"status": "COMPLETED"
}
九十一、Execution Receipt 不应该取代 Coordination Receipt
Execution Receipt 回答 what happened? Coordination Receipt 回答 why was this actor the one allowed to do it now? 两个问题不同。
九十二、完整 Multi-Agent Production Architecture
INTENT
│
▼
GOVERNANCE ENGINE
│
▼
DECISION
│
▼
AUTHORITY GRAPH
│
┌──────────┼──────────┐
│ │ │
▼ ▼ ▼
OWNER DELEGATION PRIORITY
│ │ │
└──────────┼──────────┘
▼
COORDINATION ENGINE
│
┌────────────┼────────────┐
│ │ │
▼ ▼ ▼
CONFLICT VERSION LEASE
DETECTION CHECK MANAGER
│ │ │
└────────────┼────────────┘
▼
FENCING
│
▼
CAPABILITY CHECK
│
▼
ACTION GATEWAY
│
▼
COMMAND BUS
│
▼
EXECUTION
│
▼
VERIFICATION
│
▼
COORDINATION RECEIPT
│
▼
RELEASE LEASE
│
▼
DESIRED STATE
│
▼
DRIFT DETECTION
九十三、到这里 Authority、Ownership、Coordination 三层终于分开
AUTHORITY
=
May you do it?
OWNERSHIP
=
Who governs this state?
COORDINATION
=
Who may act now?
九十四、这是多 Agent 系统非常重要的三分法
因为一个主体可以 have authority,但 not own the resource;也可能 own the resource,但 not hold the current lease。
九十五、因此新的完整判断变成
Identity valid?
↓
Authority valid?
↓
Ownership compatible?
↓
Capability valid?
↓
Coordination available?
↓
Lease current?
↓
Version current?
↓
Action allowed?
只有全部成立:EXECUTE。
九十六、把它应用到 SEO 自动化
假设 SEO Agent 准备修改 canonical。
系统读取:
Owner:
SEO Control Plane
Authority:
UPDATE_CANONICAL
Current Lease:
none
Current Version:
331
于是:
Acquire Lease
↓
Read Version 331
↓
Write
↓
Verify Canonical
↓
Commit Version 332
↓
Release
九十七、如果 Human 同时修改 canonical
Human 请求到达时 Lease Held,系统返回 RESOURCE_COORDINATED,而不是让两个写入竞争。
九十八、如果 Human 修改的是 Body
Conflict Matrix 判断 CANONICAL vs BODY = COMPATIBLE,则可以 parallel,前提是 Adapter 支持安全的 Field-level Update。
九十九、这才是真正的并发
成熟系统不是 lock everything,而是 serialize only what must be serialized。
一百、Multi-Agent Architecture 的目标不是减少 Agent
而是 allow more agents without increasing uncontrolled side effects。
一百零一、这也是为什么“Agent Swarm”概念本身没有意义
如果只有 many agents,却没有 Authority、Ownership、Coordination、Conflict Resolution,那么所谓 Swarm 很可能只是 Concurrent Automation Chaos。
一百零二、多 Agent 的成熟度不应该用 Agent 数量衡量
真正应该看的是:How safely can independent agents share production state?
一百零三、第三十四篇新的 Safety Invariants
此前我们已经有:
No Valid Decision
→
No Capability
No Capability
→
No Production Mutation
No Release Gate
→
No Exposure Expansion
No Baseline Verification
→
No Change Closure
Unresolved Critical Drift
→
No New High-risk Release
现在增加:
No Resolved Ownership
→
No Managed Write
一百零四、再增加
No Valid Coordination Right
→
No Exclusive Write
一百零五、再增加
Stale Lease Epoch
→
No Production Mutation
一百零六、再增加
Stale Resource Version
→
No Blind Overwrite
一百零七、Delegation 层增加
Delegated Authority
must never exceed
Parent Authority.
一百零八、Ownership 层增加
One Managed Field
→
One Authoritative Owner
这并不表示只有一个 Writer 实例,而表示同一个 Managed State 不能同时存在多个相互矛盾的最终权威来源。
一百零九、最后增加
No Coordination Receipt
→
No Complete Multi-Agent Write
结语:多 Agent 的难点从来不是“让更多 Agent 一起工作”,而是“让它们不会同时认为自己拥有生产控制权”
从第(二十五)篇开始,我们一直在解决一个逐步升级的问题。
第(二十五)篇:Can this action be allowed?
第(二十六)篇:Do we have enough evidence and risk context?
第(二十七)篇:Can Governance actually enforce the decision?
第(二十八)篇:Can authorized execution survive retry and partial failure?
第(二十九)篇:Can abnormal execution be reconciled?
第(三十)篇:Can systemic failure be stopped?
第(三十一)篇:Can the system learn permanently?
第(三十二)篇:Can production exposure grow only with evidence?
第(三十三)篇:Can production continuously remain equal to approved state?
到了第三十四篇:
Can many authorized actors
share production safely?
这个问题比 Can Agent A call API X? 复杂得多。
因为真正的生产环境最终一定会存在:
Agent
Agent
Agent
Human
Controller
Plugin
Scheduler
External System
而一个成熟的平台不能依赖 everyone behaves nicely,也不能依赖 hopefully they do not run at the same time。
它必须把并发治理真正做成系统能力:
Identity
↓
Authority
↓
Ownership
↓
Coordination
↓
Lease
↓
Version
↓
Fencing
↓
Conflict Resolution
↓
Execution
↓
Verification
所以第三十四篇最重要的原则可以压缩成一句话:
Authorization tells us who may act; coordination determines who may act now.
也就是:
Permission
does not create
exclusive control.
Ownership
does not create
temporary control.
Lease
does not create
authorization.
三者必须同时存在,又必须彼此独立。
最终,一个成熟的 SEO / GEO Multi-Agent Platform 应该能够在每一次 Production Mutation 前回答:
Who is the caller?
What authority
does it actually possess?
Who owns this resource?
Who owns this field?
Is authority delegated?
Where did delegation originate?
Is the delegation still valid?
Is another actor
currently coordinating
this resource?
What is the current lease epoch?
What is the current
resource version?
Does this operation conflict
with another operation?
Can the conflict be merged?
Must it wait?
May one actor preempt another?
If authority changes,
will stale actors be fenced out?
After execution,
has coordination ownership
been safely released?
只有当这些问题能够被机器化回答以后,所谓 Multi-Agent System 才不再是 multiple independent scripts calling the same APIs,而真正变成:
Governed Multi-Agent Production System
到这里,我们可以把目前形成的原则继续扩展为:
Execute Fast
Stop Faster
Recover Carefully
Learn Permanently
Release Progressively
Reconcile Continuously
Coordinate Explicitly
而这一篇最核心的新原则则是:
Shared production requires explicit coordination.
共享生产状态必须拥有明确的协调协议。
否则 more agents 并不会带来 more autonomy,而只会带来 more race conditions、more lost updates、more authority ambiguity、more unpredictable side effects。
真正成熟的 Agent 平台不是让每一个 Agent 都拥有更大的权限,而是让越来越多的 Agent 可以在:
clearly scoped authority
+
explicit ownership
+
bounded delegation
+
versioned state
+
lease-based coordination
+
enforced fencing
之下安全地共享生产系统。
这才是从 Single Agent Automation 走向 Multi-Agent Production Governance 真正需要跨越的一层。
官方依据与工程边界说明
本文提出的 Authority Graph、Coordination Control Plane、Coordination Key Registry、Conflict Domain、Conflict Matrix、Ownership Registry、Fencing Epoch、Authority Epoch、Coordination Receipt、Authority Freshness Primitive 等,是本 SEO / GEO 自动化项目为了统一 Governance、Capability、Desired State 与多 Agent 并发控制所建立的工程抽象,并不是 Kubernetes、PostgreSQL、SPIFFE 或 NIST 共同发布的一套统一 Multi-Agent Coordination 标准。
Kubernetes Lease / Leader Election:Kubernetes 使用 coordination.k8s.io 中的 Lease 对象进行节点心跳和组件 Leader Election。Lease 可以记录 holder、renew time 与 duration;协调式 Leader Election 结合 resourceVersion 的乐观并发控制,让并发候选者中只有一个更新成功。本文借鉴这一模式设计 Agent Lease、Lease Expiry 与 Coordination,但并不把 Kubernetes Lease 直接等同于 Agent Resource Ownership。官方资料:https://kubernetes.io/docs/concepts/cluster-administration/coordinated-leader-election/
PostgreSQL Explicit / Advisory Locking:PostgreSQL 提供多种显式锁以及 application-defined advisory locks。Advisory Lock 的核心限制之一在于其使用需要应用共同遵守,数据库并不会自动赋予任意业务语义。本文据此强调:Lock Registry 本身不够,真正的 Production Mutation 必须由 Action Gateway 强制验证 Coordination State。官方资料:https://www.postgresql.org/docs/17/explicit-locking.html
SPIFFE Workload Identity:SPIFFE Workload API 为工作负载提供可移植、可验证的密码学身份,使生产授权能够面向真实 workload,而不仅仅依赖应用声明的字符串名称。本文借鉴这一能力,将 Agent Instance Identity 与 Authority Graph、Capability、Lease Holder 进行关联。官方资料:https://spiffe.io/docs/latest/spiffe-specs/spiffe_workload_api/
SPIRE Delegated Identity:SPIRE Delegated Identity API 表明,被授权的 Delegate 能够代表其他 workload 获取身份,因此 delegated workload 需要被视为高信任主体。本文借鉴这一风险模型建立 Agent Delegation Chain、Authority Attenuation 与 Revocation Propagation;这些具体数据结构属于本项目设计。官方资料:https://spiffe.io/docs/latest/deploying/spire_agent/
NIST Zero Trust Architecture:NIST SP 800-207 将资源访问建立在 Subject、Resource、Policy Decision 与 Policy Enforcement 之上,强调动态、细粒度授权以及最小权限。本篇借鉴这一思想,把 Multi-Agent Authority 建立为具体 Request / Action / Resource / Environment 范围内的条件性授权,而不是把“可信 Agent”视为长期无限生产权限。官方资料:https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-207.pdf
第(三十五)篇自然承接方向
第三十四篇已经回答:
Who owns it?
Who may act?
Who may act now?
How are conflicts prevented?
下一层最自然的问题变成:
Where does an Agent's
production identity come from?
How is that identity attested?
How does an Agent obtain
short-lived credentials?
How are secrets eliminated
from Agent runtime?
How is delegated identity
constrained?
How can every production action
be cryptographically tied
to the workload that performed it?
因此第(三十五)篇可以正式进入:
Workload Identity / Attestation / Secretless Agent / Credential Broker / Delegation Chain / Identity Federation / Non-repudiation
把目前的 Authority Graph 继续向下落到真正的:
Cryptographic Workload Identity
↓
Attestation
↓
Short-lived Credential
↓
Capability
↓
Production Action
↓
Verifiable Execution Identity
从而解决 Agent 生产体系中另一个核心问题:
我们不仅要知道“哪个 Agent 名称执行了操作”,还必须能够证明:到底是哪一个经过认证的运行实例,在什么身份和授权链下完成了这次生产操作。