Guide

Workflow

How Gem Team turns AI coding into a structured, reliable engineering process.

The Orchestrator runs five phases: Init -> Route -> Plan -> Execute -> Output. Planned work uses ordered waves; standalone research uses a direct, read-only fast path.

The Orchestrator never performs specialist work itself. Implementation, debugging, testing, docs, devops, and research execution are always delegated to their owning agents; the fast path skips planning and review overhead, never delegation. The Orchestrator acts directly only to classify, route, synthesize results, ask the user, and report status.

Phase 0: Init & Clarify

Phase 0 performs a shallow, non-delegable classification:

  1. Load .gem-team.yaml, when present, and only continuity memory.
  2. Assign or validate the exact plan_id and normalize the request state (new_task, continue_plan, or extend) and intent (execute, debug, research, discuss, or challenge).
  3. Match risk signals from the supplied evidence. Do not inspect the repository or delegate to improve confidence.
  4. Assign provisional complexity once: TRIVIAL, LOW, MEDIUM, or HIGH.
  5. Ask only about a decision blocker; otherwise record one bounded assumption and route immediately.

Plan identity and state isolation

Plan selection is explicit and exact. A new task never inherits another plan's workflow state.

  • new_task -> generate a new plan_id; persistent MEDIUM/HIGH plans use it for state.
  • continue_plan or extend -> require an exact explicit plan_id; never infer one.
  • Missing or invalid IDs -> block and request the ID.
  • Every path has an ID. TRIVIAL/LOW uses it for correlation only and never reads plan artifacts.

Phase 1: Route

Routes based on plan state:

StateAction
discussAnswer directly in Phase 4
challengeRun a read-only critic review, then Phase 4
researchCall gem-researcher, then Phase 4
continue_plan with execution-only feedbackPhase 3
continue_plan with scope, wave, or criteria editsPhase 2
new_task or valid extendPhase 2

Standalone research uses its ID for correlation but creates no persistent plan, workflow state, or wave. The Orchestrator validates the research deliverable and reports it directly. Stable findings (stable: true, confidence >= 0.95) are promoted to repo memory on success. next_action: plan_follow_up promotes the work into Phase 2 under the same ID; next_action: needs_input returns the questions without continuing.

Fast-path work auto-promotes when continued execution needs dependencies, shared state, contract/risk changes, or durable evidence. Keep plan_id, preserve the current task owner and original task wave, keep completed work in its existing position, create docs/plan/{plan_id}/plan.yaml, preserve valid work/acceptance criteria, place dependent new tasks in later waves, and route only newly discovered scope through the Planner.

Phase 2: Planning

Planning depth matches complexity:

  • TRIVIAL/LOW - a clear, bounded, low-risk task uses the direct specialist path. Other work uses an in-memory wave plan. Bug fixes go to gem-debugger before gem-implementer.
  • MEDIUM/HIGH - generate or reuse the exact persistent ID and delegate to gem-planner. The Planner may promote complexity, never downgrade it.
  • Task coalescing - the Planner merges tasks that share a context base (same primary agent, overlapping files or one shared research base, no ordering dependency) and gives each cluster a shared_context block consumed by all its tasks. A merged task carries every member's acceptance criteria, so coarser tasks do not mean coarser verification.
  • Bounded planner exploration - the Planner confirms owners, boundaries, and wave placement, then stops. Deep exploration becomes a planned gem-researcher task with an explicit exploration_mode, so research cost is budgeted in the plan instead of spent silently inside planning.
  • Pre-execution review - the Planner runs a mechanical self-check before writing the plan, covering assignable owners, positive waves, per-task criteria, resolvable dependencies, and cluster consistency. A failed item is fixed and re-run once; a persistent failure returns needs_revision. Routine MEDIUM work passes on that alone; HIGH, risky, explicitly requested, or poorly evidenced work still routes to the Reviewer.

Reviewer and critic routing

Reviewer routing uses three independent axes:

  • review_mode: standard, high, or critic controls intensity.
  • review_target: plan, task, code, decision, docs, config, or integration selects the artifact.
  • review_scope: changed, affected, or full limits how much supplied evidence is examined. It is not permission to gather more: plan/task/decision reviews evaluate what they were handed, while code/config/integration reviews read the target as their subject.

Discussion bypasses delegation. Only a requested evaluation or decision routes as challenge, with critic mode, decision target, and full scope. Critic mode is read-only: it does not implement proposals, mutate files, or claim completion. Critic fields live under the reviewer handoff:

review_mode: critic
review_target: decision
review_scope: full
handoff:
  critic_subject:
    objective: str
    proposal: str
    constraints:
      - str
    alternatives:
      - str
    evidence:
      - str
    decision_needed: str
  critic_context:
    audience: str
    time_horizon: str
    success_criteria:
      - str
    known_unknowns:
      - str

Validation failures trigger a bounded replan or escalation when the budget is exhausted, progress is not measurable, or the objective or criteria would change.

Replan guardrails

Replans revise execution details, not the user's goal:

  • Plans capture an immutable baseline objective and criteria.
  • Replans use lineage (revision, replan_count) and default to a maximum of two.
  • Each replan records its reason, task changes, preserved criteria, new risks, and measurable progress.
  • The Planner may revise waves, routing, handoffs, and validation, but cannot change the baseline or decide whether the budget is exhausted.
  • Empty or repetitive replans escalate instead of looping.
  • Changed waves, handoffs, or criteria invalidate affected results and stale snapshots.
  • needs_revision revises a plan or context; needs_retry retries execution within its cap.
  • Planner revision is allowed once through planner_revision_used; reviewer findings go to the owning specialist, while non-plan reviews do not retry automatically.

Phase 3: Execution

All planned and executable work uses one wave loop. Persistent paths use docs/plan/{plan_id}/; ephemeral paths use the ID for correlation only. Standalone research is read-only until promoted to planning.

  1. Collect - load docs/plan/{plan_id}/plan.yaml on resume and on first execution alike; it is the authoritative task source, never reconstructed from the conversation. Then load the lowest incomplete wave; never skip a lower wave.
  2. Schedule - run eligible tasks in plan order, up to orchestrator.max_concurrent_agents. Retries count toward the cap. Run a cluster's tasks contiguously: sequential within a cluster, parallel across clusters.
  3. Protect ownership - tasks with overlapping ownership never run in parallel. This is also why coalescing them into one task costs no wall-clock time.
  4. Delegate - send each task only to its assigned specialist with plan_id, config_snapshot, its cluster's shared_context verbatim, task_id, task_definition, and nested handoff, in cache-lifetime order.
  5. Gate - accept each task's reported result and aggregate criteria. Invoke gem-reviewer only for HIGH/risky work, explicit requests, critic signals, or insufficient or contradictory evidence.
  6. Persist - update in-memory state or plan.yaml, initialize and track retries_used, and relay compact evidence to later waves. When a producing task returns findings for a cluster it shares, fill that cluster's shared_context verbatim and freeze it for the rest of the cluster.
  7. Retain evidence - keep detailed logs and reports in task-scoped paths. Promote reusable learning after success when confidence is at least 0.95; package reusable workflows in .apm/skills/.

Persistent plans are resumable only by exact plan ID. Ephemeral workflows never read plan artifacts.

Wave Example

Wave 1: [Debugger] diagnose error
                   ↓
Wave 2: [Implementer] fix + [Tester] verify
                   ↓
Final gate: [Reviewer] only when integration risk is triggered

Phase 4: Output

Return a structured report with:

  • Plan ID and objective
  • Task progress: completed/total and percentage
  • Wave status, blocked tasks with reasons, and next steps

Agent output policy

Agent results use the smallest role-specific schema that preserves workflow state and required handoffs. Do not repeat input fields, criteria, workflow instructions, or unused analysis. Omit empty optional fields. Keep detailed evidence in task-scoped artifacts and return only a path or manifest. There is no universal learn or per-criterion acceptance-result field.