Guide

Core Concepts

How Gem Team turns AI coding into a reliable engineering process through scoped handoffs and memory tiers.

How It Works

The delivery pipeline classifies, plans, delegates, verifies, debugs, and learns. Work runs in ordered waves; high-confidence learnings are retained.

Knowledge Layers

Gem Team organizes knowledge across six layers, each with a specific purpose:

LayerLocationPurpose
PRDdocs/PRD.yamlRequirements, decisions, goals, metrics, risks, and priorities
AGENTS.mdAGENTS.mdStable conventions, constitutional rules, and agent instructions
Plan artifactsdocs/plan/{plan_id}/Resumable wave plans, status, outputs, and task evidence
MemoryMemory tool or configured backendDurable facts, decisions, gotchas, patterns, and failure modes
Skills.apm/skills/Reusable SKILL.md procedures and domain guidance
Derived docsdocs/knowledge/Reference notes, external docs, summaries, and research outputs

Every workflow receives a plan_id. Persistent plan artifacts store workflow state, not project knowledge. Ephemeral paths use the ID for correlation only and do not access docs/plan/. AGENTS.md and repo memory hold stable knowledge-not task status or temporary assumptions. New plans never inherit another plan automatically.

Scoped Handoffs

Task scope lives in task_definition.handoff. The Orchestrator gives the Planner a bounded planning_context containing the objective, criteria, risks, and relevant discovery. The Planner returns plan.yaml; the Orchestrator derives task handoffs with constraints, target files, prior-wave outputs, and runtime evidence. Reviewers receive a dedicated review handoff, and every delegate receives a sanitized, role-scoped config_snapshot.

Plans that share a context base are coalesced before execution. Tasks with the same primary agent, overlapping files or one shared research base, and no ordering dependency become one task carrying all their acceptance criteria, with findings used by more than one task promoted into a cluster-level shared_context block. Overlapping ownership already prevents parallel execution, so merging costs no wall-clock time and removes an invocation plus a duplicated context copy.

Payload Stability

Providers cache prompts by exact prefix, so both what the workflow sends and the order it sends it in affect cost. The Orchestrator serializes every delegation payload by cache lifetime: plan-constant fields first, then cluster-shared, then per-task, with per-attempt fields last. The longest possible prefix is shared between tasks, and a retry diverges only in the tail.

Every field reused across delegations stays byte-identical: plan_id and config_snapshot per plan and agent, cluster shared_context per cluster. Nothing is subsetted, reordered, or restated, because identical bytes are what let the next call reuse a cached prefix. On resume, task scope comes from the persisted plan.yaml verbatim rather than being reconstructed from the conversation.

Three-Tier Memory

TierScopePersistence
RepoWorkspace-scopedStored with the repository
SessionConversation-scopedCleared after conversation ends
GlobalUser-scopedPersists across all workspaces

Only stable, reusable learnings with confidence of at least 0.95 are promoted after final success. Task-local evidence remains in session or workflow state. Memory entries carry a 30-day TTL to prevent stale accumulation. The Orchestrator pre-filters relevant repo memory into planning context; all agents can check repo and session memory for directly relevant findings before exploring.

Execution Model

Planned work uses one wave loop: each wave completes before the next begins, and handoffs carry evidence forward. A single bounded, low-risk task uses the direct specialist fast path. MEDIUM/HIGH work uses a persistent planner-confirmed plan. The Planner runs a mechanical self-check (owners, waves, dependency resolution, criteria) before emitting a plan, so the plan is structurally valid by construction; routine MEDIUM plans need nothing further. Judgment checks stay independent: review remains required for HIGH complexity, high-risk or critic signals, explicit requests, or insufficient or contradictory evidence. A passing self-check never replaces review.

Verification Boundary

The Orchestrator never re-verifies, re-tests, or second-guesses completed specialist work. Verification is owned exclusively by the specialist responsible for the task or plan. When a wave or plan completes, the Orchestrator accepts results as reported and moves forward. Workflow-state bookkeeping (plan status, staleness) stays with the Orchestrator: it reads state without re-running work.

See Workflow -> for the full pipeline.