Workflow
The Orchestrator runs five phases: Init -> Route -> Plan -> Execute -> Output. Planned work uses ordered waves; standalone research uses a direct, read-only fast path.
The Orchestrator never performs specialist work itself. Implementation, debugging, testing, docs, devops, and research execution are always delegated to their owning agents; the fast path skips planning and review overhead, never delegation. The Orchestrator acts directly only to classify, route, synthesize results, ask the user, and report status.
Phase 0: Init & Clarify
Phase 0 performs a shallow, non-delegable classification:
- Load
.gem-team.yaml, when present, and only continuity memory. - Assign or validate the exact
plan_idand normalize the request state (new_task,continue_plan, orextend) and intent (execute,debug,research,discuss, orchallenge). - Match risk signals from the supplied evidence. Do not inspect the repository or delegate to improve confidence.
- Assign provisional complexity once: TRIVIAL, LOW, MEDIUM, or HIGH.
- Ask only about a decision blocker; otherwise record one bounded assumption and route immediately.
Plan identity and state isolation
Plan selection is explicit and exact. A new task never inherits another plan's workflow state.
new_task-> generate a newplan_id; persistent MEDIUM/HIGH plans use it for state.continue_planorextend-> require an exact explicitplan_id; never infer one.- Missing or invalid IDs -> block and request the ID.
- Every path has an ID. TRIVIAL/LOW uses it for correlation only and never reads plan artifacts.
Phase 1: Route
Routes based on plan state:
| State | Action |
|---|---|
discuss | Answer directly in Phase 4 |
challenge | Run a read-only critic review, then Phase 4 |
research | Call gem-researcher, then Phase 4 |
continue_plan with execution-only feedback | Phase 3 |
continue_plan with scope, wave, or criteria edits | Phase 2 |
new_task or valid extend | Phase 2 |
Standalone research uses its ID for correlation but creates no persistent plan, workflow state, or wave. The Orchestrator validates the research deliverable and reports it directly. Stable findings (stable: true, confidence >= 0.95) are promoted to repo memory on success. next_action: plan_follow_up promotes the work into Phase 2 under the same ID; next_action: needs_input returns the questions without continuing.
Fast-path work auto-promotes when continued execution needs dependencies, shared state, contract/risk changes, or durable evidence. Keep
plan_id, preserve the current task owner and original task wave, keep completed work in its existing position, createdocs/plan/{plan_id}/plan.yaml, preserve valid work/acceptance criteria, place dependent new tasks in later waves, and route only newly discovered scope through the Planner.
Phase 2: Planning
Planning depth matches complexity:
- TRIVIAL/LOW - a clear, bounded, low-risk task uses the direct specialist path. Other work uses an in-memory wave plan. Bug fixes go to
gem-debuggerbeforegem-implementer. - MEDIUM/HIGH - generate or reuse the exact persistent ID and delegate to
gem-planner. The Planner may promote complexity, never downgrade it. - Task coalescing - the Planner merges tasks that share a context base (same primary agent, overlapping files or one shared research base, no ordering dependency) and gives each cluster a
shared_contextblock consumed by all its tasks. A merged task carries every member's acceptance criteria, so coarser tasks do not mean coarser verification. - Bounded planner exploration - the Planner confirms owners, boundaries, and wave placement, then stops. Deep exploration becomes a planned
gem-researchertask with an explicitexploration_mode, so research cost is budgeted in the plan instead of spent silently inside planning. - Pre-execution review - the Planner runs a mechanical self-check before writing the plan, covering assignable owners, positive waves, per-task criteria, resolvable dependencies, and cluster consistency. A failed item is fixed and re-run once; a persistent failure returns
needs_revision. Routine MEDIUM work passes on that alone; HIGH, risky, explicitly requested, or poorly evidenced work still routes to the Reviewer.
Reviewer and critic routing
Reviewer routing uses three independent axes:
review_mode:standard,high, orcriticcontrols intensity.review_target:plan,task,code,decision,docs,config, orintegrationselects the artifact.review_scope:changed,affected, orfulllimits how much supplied evidence is examined. It is not permission to gather more:plan/task/decisionreviews evaluate what they were handed, whilecode/config/integrationreviews read the target as their subject.
Discussion bypasses delegation. Only a requested evaluation or decision routes
as challenge, with critic mode, decision target, and full scope. Critic mode
is read-only: it does not implement proposals, mutate files, or claim
completion. Critic fields live under the reviewer handoff:
review_mode: critic
review_target: decision
review_scope: full
handoff:
critic_subject:
objective: str
proposal: str
constraints:
- str
alternatives:
- str
evidence:
- str
decision_needed: str
critic_context:
audience: str
time_horizon: str
success_criteria:
- str
known_unknowns:
- str
Validation failures trigger a bounded replan or escalation when the budget is exhausted, progress is not measurable, or the objective or criteria would change.
Replan guardrails
Replans revise execution details, not the user's goal:
- Plans capture an immutable baseline objective and criteria.
- Replans use lineage (
revision,replan_count) and default to a maximum of two. - Each replan records its reason, task changes, preserved criteria, new risks, and measurable progress.
- The Planner may revise waves, routing, handoffs, and validation, but cannot change the baseline or decide whether the budget is exhausted.
- Empty or repetitive replans escalate instead of looping.
- Changed waves, handoffs, or criteria invalidate affected results and stale snapshots.
needs_revisionrevises a plan or context;needs_retryretries execution within its cap.- Planner revision is allowed once through
planner_revision_used; reviewer findings go to the owning specialist, while non-plan reviews do not retry automatically.
Phase 3: Execution
All planned and executable work uses one wave loop. Persistent paths use docs/plan/{plan_id}/; ephemeral paths use the ID for correlation only. Standalone research is read-only until promoted to planning.
- Collect - load
docs/plan/{plan_id}/plan.yamlon resume and on first execution alike; it is the authoritative task source, never reconstructed from the conversation. Then load the lowest incomplete wave; never skip a lower wave. - Schedule - run eligible tasks in plan order, up to
orchestrator.max_concurrent_agents. Retries count toward the cap. Run a cluster's tasks contiguously: sequential within a cluster, parallel across clusters. - Protect ownership - tasks with overlapping ownership never run in parallel. This is also why coalescing them into one task costs no wall-clock time.
- Delegate - send each task only to its assigned specialist with
plan_id,config_snapshot, its cluster'sshared_contextverbatim,task_id,task_definition, and nestedhandoff, in cache-lifetime order. - Gate - accept each task's reported result and aggregate criteria. Invoke
gem-revieweronly for HIGH/risky work, explicit requests, critic signals, or insufficient or contradictory evidence. - Persist - update in-memory state or
plan.yaml, initialize and trackretries_used, and relay compact evidence to later waves. When a producing task returns findings for a cluster it shares, fill that cluster'sshared_contextverbatim and freeze it for the rest of the cluster. - Retain evidence - keep detailed logs and reports in task-scoped paths. Promote reusable learning after success when confidence is at least 0.95; package reusable workflows in
.apm/skills/.
Persistent plans are resumable only by exact plan ID. Ephemeral workflows never read plan artifacts.
Wave Example
Wave 1: [Debugger] diagnose error
↓
Wave 2: [Implementer] fix + [Tester] verify
↓
Final gate: [Reviewer] only when integration risk is triggered
Phase 4: Output
Return a structured report with:
- Plan ID and objective
- Task progress: completed/total and percentage
- Wave status, blocked tasks with reasons, and next steps
Agent output policy
Agent results use the smallest role-specific schema that preserves workflow state and required handoffs. Do not repeat input fields, criteria, workflow instructions, or unused analysis. Omit empty optional fields. Keep detailed evidence in task-scoped artifacts and return only a path or manifest. There is no universal learn or per-criterion acceptance-result field.