This commit is contained in:
2026-08-24 08:11:59 -07:00
commit 53df0eed10
275 changed files with 133056 additions and 0 deletions
+237
View File
@@ -0,0 +1,237 @@
# Optimization and Performance Policy
Codex MMO optimizes for completed, validated work—not maximum agent count. A model or worker is useful only when it improves at least one of these dimensions without imposing greater integration cost:
- Wall-clock latency through genuine parallelism.
- Capability through specialization or stronger reasoning.
- Context isolation for noisy or independent branches.
- Confidence through independent verification.
- Access-route cost or local-resource efficiency.
- Operational control through durable, cancellable jobs.
## The delegation decision
The root should evaluate five variables together:
1. **Difficulty** — how much reasoning or implementation capability is required?
2. **Independence** — can the branch proceed without blocking or duplicating root work?
3. **Criticality** — is it on the immediate critical path?
4. **Verifiability** — how cheaply can the result be checked from primary evidence?
5. **Execution economics** — launch overhead, provider cost, latency, resource pressure, and integration burden.
A strong default decision tree is:
```text
Is the task an immediate critical-path blocker?
├── yes: root usually handles it directly
└── no
├── difficult/high-consequence and independent: flagship peer
├── substantive bounded engineering: implementation specialist
├── cheap, literal, easily verified evidence: low-cost scout
└── no meaningful benefit from delegation: root handles it
```
Difficulty alone does not determine delegation. A hard independent review is an excellent flagship-worker task; a hard architectural decision that blocks all progress normally belongs in the root thread.
## Root productivity
The root is an active engineer and integrator. Its generated policy requires:
1. Identify the immediate critical path.
2. Identify independent side work.
3. Spawn only useful branches.
4. Continue non-overlapping root work immediately.
5. Consume results as they become relevant.
6. Wait only at a real dependency barrier.
7. Reconcile contradictions from primary evidence.
8. Review material changes and run integrated validation.
The runtime records explicit Agent MCP wait events and root events emitted during worker intervals. Evaluation reports wait seconds/ratio, event coverage, worker overlap, and time to first useful result; event coverage is not proof of continuous cognition. A profile that causes the root to spawn and immediately idle is still misconfigured even if worker utilization looks high.
## Native versus MCP execution
Native Codex agents and Agent MCP solve different performance problems.
Use native agents when:
- Launch latency matters.
- The task is small or read-heavy.
- Tight `/agent` interaction is useful.
- Built-in Codex authentication and native behavior should remain direct.
- External scope/contract enforcement is not essential.
Use MCP when:
- Exact model/route pinning must be auditable.
- The worker is low-trust or inconsistent.
- The task is long-running or asynchronous.
- Cancellation, durable results, or operator inspection matter.
- Nested delegation needs enforceable lineage and budgets.
- Workspace writes need isolated worktrees, explicit scopes, and audited patch integration.
- Structured output must be validated.
Hybrid profiles use native first-level branches for lower latency and MCP for strict or nested work. This avoids paying process-launch overhead universally without giving up enforcement where it has value.
## Weak-model economics
A cheap model is not automatically efficient. Its net value is:
```text
saved strong-model work
- launch and prompt cost
- verification cost
- contradiction-resolution cost
- rework caused by incorrect output
```
A weak or conflict-prone model should therefore receive only tasks that are:
- Narrow.
- Low ambiguity.
- Read-only.
- Mechanically checkable.
- Cheap to repeat or discard.
- Unlikely to contaminate architecture or final decisions.
The bundled Qwen role is restricted to evidence gathering for precisely this reason. The same trust policy can be applied to any model. Low-trust output is evidence to verify, never authority.
Do not delegate a command the root can execute directly in less time than starting and checking a worker.
## Parallelism policy
Concurrency ceilings are limits, not targets.
Read-heavy work can be parallelized aggressively when branches are independent:
- Repository mapping.
- Independent diagnosis.
- Specification extraction.
- Review from different perspectives.
- Test and log analysis.
Write-heavy work requires stronger partitioning:
- Assign disjoint files or components.
- Require MCP write scopes and isolated Git worktrees where enforceability matters.
- Avoid logically coupled edits even when paths differ.
- Keep one root responsible for reading, accepting/rejecting, explicitly integrating, inspecting, and testing patches.
- Cancel duplicate branches once one result makes the other unnecessary.
Too many agents increase context duplication, provider queuing, file conflicts, and synthesis time. Start with two or three useful branches; increase only when evaluation shows a wall-clock or quality benefit.
## Resource groups
Resource groups model shared capacity rather than model names. Examples include:
- A ChatGPT subscription.
- An OpenCode Go account.
- A Z.AI plan.
- A direct API rate pool.
- One local GPU.
Each active MCP role consumes `resource_units`. Admission is atomic across sessions. This prevents a six-agent profile from launching six copies against a one-slot local model server or exhausting one provider while other capacity sits idle.
Tune group capacity in a user catalog fragment based on measured provider and hardware behavior. Do not encode a physical capacity assumption in profile instructions.
## Context isolation
Worker prompts include the delegated objective, role instructions, output contract, workspace context, and parent lineage—not the entire root transcript by default. This reduces token duplication and keeps independent branches independent.
Good briefs state:
- Exact scope.
- Expected artifact or decision.
- Relevant constraints.
- Required evidence and validation.
- What the worker must not decide.
Bad briefs paste the complete conversation, ask an agent to “solve everything,” or omit success criteria.
## Structured results
Strict contracts reduce synthesis cost by making important outputs predictable. They are particularly valuable for:
- Low-trust evidence collectors.
- Review findings.
- Implementation summaries.
- Proposal comparisons.
- Test reports.
Keep contracts small. Requiring a large nested schema can consume more model effort than it saves. Full logs and diffs remain as artifacts; the structured result should contain the decision-relevant subset.
## Contradiction handling
Never use model voting as a substitute for verification. When agents disagree:
1. Identify the exact conflicting propositions.
2. Locate source code, specifications, tests, logs, or reproducible commands that discriminate between them.
3. Run or inspect that evidence.
4. Ask a specialist only when the evidence itself requires expert interpretation.
5. Record residual uncertainty.
A stronger model is not automatically correct; a majority is not automatically correct; an independent result is valuable only when its assumptions and evidence are visible.
## Gateway reuse
Switchyard runs per resolved transport/model route set. Sessions with equivalent gateway identity can reuse one gateway; different route semantics remain isolated. This gives:
- Stable model IDs.
- No route collisions during profile switching.
- Credential minimization.
- Reuse without globally exposing unused routes.
- Independent shutdown of idle gateways.
Profiles using only direct Codex routes avoid gateway startup entirely.
## Reasoning levels
Use the highest reasoning setting only where its marginal value exceeds latency and token cost.
Recommended defaults:
- Root or flagship architect: high to maximum supported level.
- Substantive implementation: high.
- Mechanical evidence collector: reasoning disabled or lowest reliable mode.
- Reviewer: high when subtle correctness is material.
- Plan mode: explicitly configured so the client does not silently fall back to a lower preset.
A profile binding is validated against declared model and route reasoning metadata. Labels are not assumed to be semantically identical across access products; catalog entries record the selector accepted by that exact path.
## Evaluation-driven tuning
Use matched variants within the workload-specific suite:
```bash
codex-mmo eval run --profile adaptive-engineering --suite adaptive-change
codex-mmo eval run --profile adaptive-engineering --suite adaptive-change --trial-mode release
codex-mmo eval compare RUN_ID_A RUN_ID_B
```
Compare:
- Success rate, workload metrics, and objective validation.
- Wall-clock time.
- Observed root MCP wait ratio, peak workers, worker overlap, and time to first result.
- Actual worker-role/model/route/serving routing and result-acceptance rate.
- Worker completion/failure rate.
- Result acceptance and contract failures.
- Retry and cancellation counts.
- Conflicts and out-of-scope warnings.
- Token usage plus separate actual-API, API-equivalent, subscription, and local-resource ledgers.
- Integration corrections.
Optimization changes should be justified by these results, not by intuition about model rank.
## Common anti-patterns
- Spawning every available model on every task.
- Delegating the critical path and immediately waiting.
- Giving weak models ambiguous or high-consequence work.
- Letting multiple writers edit overlapping logic.
- Using the same prompt for every role.
- Allowing unbounded nested delegation.
- Retrying a failed branch repeatedly without changing context or approach.
- Treating structured output as proof of correctness.
- Running both Codex Ultras autonomous orchestration and an external mesh without explicit accounting.
- Routing built-in Codex models through an unnecessary translation layer.
- Measuring utilization instead of completed validated work.