Guide

Optimizations

Best practices for optimizing Gem Team's agent execution, output hygiene, and scoped handoffs.

Cost and performance are core design principles. These practices keep execution and context focused.

Agent Execution Optimization

See Agents -> for model routing recommendations.

Output Hygiene

Agents reduce token waste through:

  • Native flags - use --quiet, --oneline, and maxResults to truncate output at the source.
  • Pipeline truncation - use head, tail, or grep only when native flags are insufficient.
  • Context pruning - send execution scope through task_definition.handoff, bounded planning_context to the Planner, a dedicated review handoff to the Reviewer, and a sanitized config_snapshot to every delegate.
  • Proportional architecture - use YAGNI/KISS to choose the smallest architecture, specialist set, wave sequence, and validation path that satisfies the baseline.
  • Minimal results - return only workflow state, failure classification, role-specific evidence, and downstream handoffs.
  • Evidence by reference - keep logs, traces, screenshots, and reports in task artifacts; return a path or manifest.
  • Conditional learning - return learnings only for stable, reusable, repeated, or persistent findings; promote them after success at confidence >= 0.95.

Context Efficiency & Cached Token Usage

Gem Team is designed for minimal context footprint. Each wave sends only the delta - what changed and what the next agent needs - rather than replaying the full history.

Why It Matters

MetricTypical AI ChatGem Team
Context per taskFull historyScoped handoff + delta
Re-sent tokens per wave100%~20-40% (only new evidence)
Cache hit rateLow (churn)High (stable structure)
Token cost predictabilityVolatileProportional to scope

How Context Stays Lean

  • Scoped handoffs - each agent receives only task_definition.handoff: target files, constraints, prior-wave outputs, and runtime evidence. No full repo, no full history.
  • Context coalescing - tasks that share a context base (same agent, overlapping files or one shared research base, no ordering dependency) merge into one task. They already could not run in parallel because their ownership overlaps, so merging costs no wall-clock time and saves an invocation plus a duplicate context copy. Merged tasks keep every member's acceptance criteria.
  • shared_context - findings consumed by two or more tasks are inlined once per cluster and injected verbatim into each consumer. Inlining adds tokens the single-consumer rule would have avoided, but a path costs every consumer a separate uncached tool read. When the findings come from an executed producer task, the Orchestrator fills the block verbatim from its return and freezes it for the rest of the cluster.
  • Cache-lifetime field order - every delegation payload is serialized plan-constant first, then cluster-shared, then per-task, with per-attempt fields last. Prefix matching only reaches the first divergence, so a stable field placed after a per-task one is never reused. Every reused field stays byte-identical across the calls that share it; per-task differences live only in the tail.
  • Stable resume - a resumed plan loads task scope from the persisted plan.yaml rather than rebuilding it from the conversation, so the original bytes are reused instead of a paraphrase that misses the cache.
  • Bounded planning_context - the Planner gets objective, criteria, risks, and discovery - not the entire conversation.
  • Evidence by reference - logs, traces, screenshots, and reports live in task-scoped artifacts. Agents return a path or manifest, not the content. Exception: a finding consumed by two or more tasks is inlined, because a path costs every consumer a fresh uncached tool read.
  • Compact config_snapshot - every delegate gets a sanitized, role-scoped configuration. No global config bloat.
  • Proportional architecture - YAGNI/KISS selects the smallest agent set, wave sequence, and validation path. Fewer agents = less context multiplied.

What drives cache hits

LLM providers cache prompt prefixes. Gem Team's structured handoff format means effective token cost is proportional to new work, not total conversation history. See the Cost in practice section below for real numbers.

Cost in practice

Gem Team's cache efficiency is not theoretical. Observed data from real development workloads on DeepSeek V4.1 Flash:

MetricValueNotes
Tokens processed82.8M+Across 666 agent runs
Avg. context/run~124K tokensLarge-context workflows, not tiny snippets
Typical API callsub-$0.001100K+ input, mostly cached
Typical latency2-5sFor 100K-190K input, 100-1000 output tokens

Why it's cheap

DeepSeek V4.1 Flash pricing: $0.15/M uncached input, $0.003/M cached input, $0.60/M output.

For a typical 150K-token context:

Cache stateApprox cost per call
100% miss$0.0225
95% cached~$0.00153
99% cached~$0.00067
100% hit$0.00045

Gem-Team's structured handoffs mean most of each 100K+ token call hits cache. The cost difference between a cache miss and a hit is over 30x. That is what makes sub-penny agent runs possible at scale.

What drives the cache hits

  • Stable system prompts - agent definitions and rules don't change between tasks, so they stay cached.
  • Consistent handoff shape - the task_definition structure is identical across waves, maximizing prefix reuse.
  • Cache-lifetime field order - retries and follow-up calls diverge only in the payload tail, so the scope prefix stays cached.
  • Cluster-contiguous scheduling - tasks sharing a context_cluster run back to back, so the shared prefix stays warm inside the provider's cache window.
  • Minimal churn - only new evidence appends to the context window; previous waves' outputs are referenced, not re-sent.

Cache hits are exact-prefix, so ordering is part of the mechanism, not a detail: reordering context invalidates it even when the content is identical.

Observed during Gem-Team development using DeepSeek V4.1 Flash via CommandCode. Actual cost depends on provider, cache state, and input/output ratio.

Handoff Optimization

Handoff scope and memory tiers are covered in Core Concepts ->.

Execution agents receive scope through each task's nested task_definition.handoff. The Planner receives only objective-aligned clarifications and sufficient discovery in planning_context, and may request more context only for a concrete blocker. Relay relevant evidence only; promote stable learnings after final success at confidence >= 0.95.

Verify the cache, don't assume it

Cache accounting belongs to the provider, not to Gem Team. Read it from the response usage fields (cache_read_input_tokens, or prompt_tokens_details.cached_tokens) for a plan with several same-cluster tasks. If your client does not expose cache counters, treat the figures above as a design intent, not a measurement.