Agent Team
Agent Definitions
Definitions live in .apm/agents/ and are routed by the Orchestrator.
Orchestrator
Coordinates the Phase 0-4 workflow, delegates work, tracks state, and enforces verification gates. It classifies from supplied evidence and never executes specialist work directly. Classifies failures into retry, fixable, replan, flaky, regression, platform-specific, and test-bug categories so each routes to the right agent. Applies relational invariant fallbacks when agent outputs miss conditional fields - inferring intent and filling safe defaults instead of rejecting valid work. Loads task scope from the persisted plan.yaml on resume, propagates each specialist's typed handoff output into dependent tasks' handoff.relevant_context, and fills cluster shared_context verbatim from producer returns so downstream payloads stay byte-identical. Promotes stable research findings to repo memory and injects relevant memory into planning context.
Core Agents
| Agent | Role |
|---|---|
| Researcher | Explores patterns, dependencies, architecture, and docs in five budgeted modes: scan, deep, audit, trace, and question. Runs directly for research or inside a persistent wave plan under a Planner-set budget. |
| Planner | Validates the baseline and builds a minimal plan with ordered waves, routing, handoffs, and risks. Persists docs/plan/{plan_id}/plan.yaml, applies YAGNI/KISS, slices along concern boundaries, coalesces tasks that share a context base, and runs a scope-reduction gate plus a mechanical self-check. Confirms owners; delegates deep exploration. |
| Implementer | Implements features, fixes, and refactors with TDD. Covers happy paths, invariants, boundaries, errors, input variation, and state transitions. Bug fixes require diagnosis first. Prefers the shortest correct diff and leaves a runnable check for non-trivial changes. Applies SOLID, fail-fast, and least-surprise. Every major decision has a one-line reason. No dead buttons. |
Quality & Review Agents
| Agent | Role |
|---|---|
| Reviewer | Reviews plans and implementations for quality, security, contracts, compliance, and regressions. Runs an over-engineering pass for code and integration targets. Provides read-only critic reviews for decisions. Verifies every decision has a reason. Flags dead buttons and external script patches as blocking issues. |
| Debugger | Reproduces failures, finds root causes, and bisects regressions. Adds a minimal reproduction test before recommending a fix; never implements it. |
| Browser Tester | Checks browser and UI flows, with conditional visual, accessibility, performance, network, and regression testing. |
| Code Simplifier | Removes dead code, reduces complexity, consolidates duplicates, and improves naming while respecting existing design intent. Every refactoring has a one-line reason. No buzzwords. Removes AI-slop comments. |
Specialized Agents
| Agent | Role |
|---|---|
| DevOps | Handles infrastructure, CI/CD, containers, production approvals, health checks, and rollback. |
| Documentation | Creates technical docs, READMEs, API docs, diagrams, walkthroughs, and PRD updates. No buzzwords. Every section exists because the product needs it. No fabricated statistics. |
| Mobile Tester | Runs scope-selected mobile E2E checks with Detox, Maestro, or Appium. |
| Skill Creator | Extracts high-confidence patterns into reusable SKILL.md files, scripts, references, and assets. |
Reviewer routing
gem-reviewer uses three axes:
review_mode:standard,high, orcriticcontrols review intensity.review_target:plan,task,code,decision,docs,config, orintegrationselects the target.review_scope:changed,affected, orfulllimits how much supplied evidence is examined. It is not permission to gather more:plan/task/decisionreviews evaluate what they were handed, whilecode/config/integrationreviews read the target as their subject.
Single bounded, low-risk work uses the direct specialist fast path. Other TRIVIAL/LOW work uses ephemeral planning without Planner or Reviewer calls.
For MEDIUM/HIGH work, the Orchestrator supplies bounded discovery in planning_context; the Planner owns waves and handoffs, while the Orchestrator owns state, retries, approvals, review, and replan limits.
MEDIUM/HIGH plans receive a mechanical self-check from the Planner before they are written (owners, waves, dependencies, criteria, cluster consistency); routine MEDIUM plans need nothing further. Explicit requests, insufficient evidence, HIGH complexity, and high-risk or critic signals require independent review.
Critic mode is read-only. It evaluates assumptions, counterexamples, risks, alternatives, reversibility, and decision blockers without changing files or claiming completion. Results include critic_verdict, challenges, alternatives, and decision_blockers, each with evidence, impact, and action. There is no standalone gem-critic agent.
Critic requests carry these objects under handoff:
review_mode: critic
review_target: decision
review_scope: full
handoff:
critic_subject:
objective: str
proposal: str
constraints:
- str
alternatives:
- str
evidence:
- str
decision_needed: str
critic_context:
audience: str
time_horizon: str
success_criteria:
- str
known_unknowns:
- str
Model Compatibility
All agent definitions use valid YAML contracts and consolidated output fields so they work across commercial and open models without prompt-engineering workarounds. The Orchestrator applies relational invariant fallbacks - when an agent output violates a conditional requirement, it infers the most likely intent and fills the gap instead of rejecting valid work.
Model Routing
The Orchestrator supports two configurable model tiers through
model_routing in .gem-team.yaml:
| Tier | Agents | Purpose |
|---|---|---|
| Premium | Planner, Debugger, Reviewer | Reasoning-heavy planning, diagnosis, challenge, and verification |
| Explore | Researcher, Implementer, Browser Tester, Mobile Tester, DevOps, Documentation, Skill Creator, Code Simplifier | Fast exploration and bounded project work |
When enabled, the Orchestrator passes the configured tier model to each delegated agent. The Orchestrator itself is not routed through these tiers.
| Agent tier | Recommended model | Use case |
|---|---|---|
| Explore | Fast/flash (gpt-5.6-luna, claude-3.5-haiku, gemini-2.5-flash) | Implementation, docs, research, and simple checks |
| Premium | Strong/pro (gpt-5.6-sol, claude-4, gemini-2.5-pro) | Planning, analysis, compliance, and high-risk verification |
Configurable in .gem-team.yaml. Swap with equivalent models from your provider.
Packaged Skills
Reusable, task-specific guidance lives separately from agent definitions in .apm/skills/:
| Skill | Purpose |
|---|---|
gem-devops-guidelines | Deployment strategies, CI/CD, containers, health checks, rollback, and production readiness. |
Skills are loaded for matching tasks and keep specialist agent definitions focused.