Delivery / An interactive field guide
One request.
Every handoff.
Follow the work as it moves through the team. Watch the files, commits, decisions, and reviews accumulate—then see exactly what comes back to you.
01 / Follow a run
Scripted example · starts paused/featureAdd a dashboard showing sync status for all customer listings
- Stage
- Entrance
- Round
- —
- $SEQ
- —
mainEntrance
A human brings the intent.
- 1 · Red
- 2 · Green
- 3 · Refactor + validate
- 4 · Commit passing task
Four ledgers / One shared history
What this beat leaves behind.
Color identifies the destination.
Scrub backward. The ledgers travel with you.
Local git
0Shared SQLite
0 activeGitHub
0 commitsA teaching model, not live telemetry. Run counts cover the depicted specialist runs; orchestrator logging is not counted here. Scores, counters, findings, commit labels, and triage judgments are illustrative. Bug identifiers remain source templates.
Optional / Across multiple runsHow the team improves itselfExplore the learning branch ↗
Delivery leaves evidence. A later, human-invoked learning pass turns recurring evidence into proposed changes. The pipeline only checks the configured cadence and mentions when analysis is due.
- 01 / Human invokes
Log Analyst
Reads accumulated logs and current agent definitions. Compares decisions, findings, outcomes, and input-quality trends across runs, then proposes specific improvements with evidence and confidence.
docs/agent-analysis/YYYY-MM-DD.mdReads SQLite · does not write to it - 02 / Human decides
Approve selectively
Review the evidence and exact proposed change. Approve, defer, or reject each proposal independently. A recurring finding is a candidate for judgment, not permission to rewrite the team.
Approved scope of changeNo approval → no policy change - 03 / Human invokes
Skill Builder
Receives the selected analysis proposal or a direct instruction and produces skill files. This is a separate invocation in the main checkout; the feature pipeline does not automatically start it.
Skill files for the selected changeBounded input → reusable instruction - 04 / Observe again
Future runs
Agents use the revised instructions on later work. Compare subsequent findings and outcomes to assess whether the change helped. Writing a skill is not evidence that it worked.
New decisions, findings, outcomesEvidence feeds the next analysis
The counts in the four ledgers above belong to the selected delivery or triage example. Opening this explanation does not add synthetic runs. Log Analyst deliberately does not log its own run to the database.
02 / The project remembers
The code tells you what exists.
The record tells you how you got here.
Each run can leave more than an implementation: the intent, choices, alternatives, review findings, and outcomes behind it. Kept with the project, those records give the next person—and the next agent—a place to begin.
SQLite connects the evidence.
Runs connect decisions, events, findings, and reflections to the work that produced them. Recorded rationale and expected versus observed outcomes make it possible to compare choices across features and agents.
Query relationships and recurring patternsMarkdown preserves the story.
Briefs, specifications, engineer reports, and reviews remain with the project. Numbered rounds preserve the progression; commits show when those artifacts changed and what was delivered.
Read the context behind the implementationScope capture preserves the opportunity.
An agent can notice a bug, missing capability, or cleanup opportunity, record it in its report, and continue its assigned task. The orchestrator later checks duplicates and files the surviving ideas.
Notice → record → continue the current taskAsk the project
Make the history useful.
Choose a question to see which evidence you would follow.
The handoff that compounds
The next run starts
with more context.
Every recorded choice gives the next contributor something to work from. They can understand the intent, question the tradeoffs, and find the unfinished work before changing the application.
Explore the evidence behind each runCarried forward from earlier work
- Intent, preserved.Markdown
Briefs and specs explain what the work was meant to achieve.
- Choices, traceable.SQLite
Decisions and outcomes expose the recorded reasons and tradeoffs.
- Opportunities, retained.Scope capture
Captured ideas keep useful future work available for prioritization.
03 / Take the mechanism with you
The language can change.
The contracts still matter.
You can build this team for another stack. Keep the boundaries explicit, make the evidence durable, and keep the final decision human.
Give judgment a boundary.
Architect, Design, Engineer, and the four reviewers run as subagents with bounded inputs and outputs. Discovery and bug triage stay live because they need a conversation with the human.
Explore the agents →Let the files carry the handoff.
A brief becomes a spec, a spec becomes a design contract, and a report becomes review input. The orchestrator checks artifacts before advancing and can resume from the first missing stage.
Explore the architecture →Separate evidence from delivery.
The worktree isolates the branch. A shared database retains agent activity across worktrees. Local commits preserve the work before the final push makes that branch visible remotely.
Explore agent logging →Scope capture / Destination matters
Notice now. File at the end.
Agents record ideas in their own reports. After acceptance, the orchestrator sweeps all rounds once and checks for duplicates before filing.
[needs-discovery]- GitHub Project card
[tech-debt]- GitHub Project card
[bug]- Repository issue
- Already exists
- Skip the duplicate
Source notes & what this example simplifies
Grounded in the Rails Agentic Engineering Team and the supplied pipeline visualization spec. Each beat links to its implementation source. See the source mapping manifest for the detailed mapping.
The default example uses two review rounds. The failure scenario uses the default three-round escalation threshold; projects can configure it. The database illustrates specialist start/end lifecycles and input_quality reflections, not a complete production log. Playback timing is editorial, not a performance estimate.
The scope example follows the spec’s explicit timeline: two Project cards and one skipped duplicate. A new bug would instead be a repository issue. After PR creation, the source commits and pushes the PR URL into the summary file; it does not prescribe editing the PR body.
Triage normally reads GitHub without changing existing issues. Reclassifying a report into a feature can lead to filing only with explicit human confirmation. The main triage track stops at its recommendation.