Two agents that do
on-call work
An alert fires at 3am. Something changed twenty minutes ago and nobody remembers what. Oncall investigates it — reads the metrics, the logs, the cluster events and the deploy history — then opens a pull request containing the revert and its own reasoning. A human reviews and merges. That merge is the approval; the agent never merges its own work.
A second agent, Automation-Engineer, turns spoken requirements into n8n workflows. It shares no tools with the first one, and it cannot reach the cluster at all.
The demo, end to end
A bad release breaks a liveness probe. Nothing here is simulated: the alert is a real Grafana rule over real Prometheus data, the cluster is a real kind cluster, and the pull request is a real pull request.
- A commit changes the demo service's health endpoint. ArgoCD auto-syncs it. The liveness probe now points at a path that 404s.
- Pods start failing their probe and restarting. kubelet · CrashLoopBackOff
- Grafana fires
ReplicasUnavailable. The rule is Terraform-managed. Its webhook posts to the orchestrator, bearer-authenticated. - The orchestrator admits the alert and starts a healing session. Deduplicated per fingerprint, rate-limited, flap-delayed. One session per incident, not one per notification.
- SRE-Oncall investigates. Grafana MCP for the alert rule and PromQL · Kubernetes MCP for pod status, events and the previous container's logs · ArgoCD MCP for "what synced, and when"
- It correlates the failure to the sync. "This started four minutes after sync abc123, which changed the probe path" — the single most valuable sentence in an incident.
- It opens a revert pull request — without asking. Creating a branch and a PR changes nothing that runs. The PR is the artefact a human reads, so producing one must not itself wait on approval.
- The PR body carries the whole case. CAUSE with evidence · CHANGE · WHY · EXPECT · RISK. The reviewer never has to read back through the session.
- A human reviews and merges. This is the gate.
The agent cannot merge —
merge_pull_requestis gated and the prompt forbids it outright. - ArgoCD syncs the revert. The probe recovers. Verified by re-querying the alert's own expression, not by a pod turning Ready — a crashlooping pod is Ready in bursts.
- The alert resolves and the agent posts the summary in the same Slack thread. Then drafts a postmortem into Notion.
Who calls what
The healing loop as it actually runs, one lane per phase. Every hop names the MCP server that carries it — that is the part worth being precise about, because it is where the agent's reach begins and ends.
The agent never merges. It reaches production only by writing a document a person reads. Opening that document is ungated on purpose — stopping to ask permission to open a PR leaves the incident burning while you wait for permission to do the one thing that helps.
Slack → run an automation
Two Slack apps, two Socket Mode connections. A mention fires in any channel a bot is in, so each bot is scoped to its own — otherwise both would answer in the same room.
Two agents, disjoint tools
The split is not organisational tidiness. The two jobs pull the harness configuration in opposite directions, and giving one agent both sets of tools makes each one worse.
| Agent | Model | Harness config | Why |
|---|---|---|---|
| sre-oncall prompts/base.md |
openai/ SRE_ONCALL_MODEL |
askUserQuestions: falseiterationLimit 200 · no sub-agents · sandbox on |
Woken by an alert with nobody watching the screen. A blocking question is not a gate, it is a hang. Its human checkpoints are deliberate and elsewhere: the Slack approval, and the pull request review. |
| automation- prompts/automation-engineer.md |
openai/ AUTOMATION_ENGINEER_MODEL |
askUserQuestions: trueiterationLimit 100 · no sub-agents · sandbox on |
Only ever runs because a person started it and is waiting. Turning "notify me when a deploy fails" into a specific workflow is impossible without asking which channel, which schedule, which credential. |
Sub-agents are off on both, deliberately. A fresh sub-agent starts cold and re-reads every tool schema already paid for. The binding constraint here is a tokens-per-minute ceiling, not reasoning depth — investigations have died at 535k and 618k input tokens. The prompt says not to delegate; the config enforces it.
Which agent holds which MCP servers
- grafana 14 gated
- kubernetes writes gated
- argocd writes gated
- terraform writes gated
- github 2 gated
- notion 1 gated
- n8n-builder 11 gated
- n8n-tools all gated
Because the sets are disjoint, the automation agent cannot touch production even if a requirement asks it to — its blast radius is the workflows it writes. The split also stops each agent paying tool-schema tokens for the other's servers on every single request.
Every tool, and where the gate is
Gating is declared per MCP attachment and enforced by the harness. The interesting part is not that gates exist — it is where they are not, because each of those is a decision that had to be argued for.
| Server | Enabled | Needs approval | The thing worth explaining |
|---|---|---|---|
| grafana | all but onegrafana_api_request disabled outright |
14 named write tools | Gated by name, not by @write. alerting_manage_rules reads
and writes behind one name — operation: "list" is how every
investigation begins, so gating by shape stopped the agent before it could look at the
alert that woke it. Left ungated; the prompt forbids mutating a rule in words. An
explicit trade, stated rather than hidden. |
| kubernetes | all | @write, @destructive |
Reading is free — pod status, events, and the previous container's logs, which is where an OOM kill or a panic is actually visible. Every mutation stops. |
| argocd | all | @write, @destructive |
A rollback is a production change. The reads are what make "started four minutes after sync abc123" sayable. |
| terraform | all | @write, @destructive |
Registry docs and plan/validate are free; apply
stops. Terraform owns the alerting config, so a wrong threshold is an HCL change with a
plan attached to the PR. |
| github | 11 named tools the server exposes 44 |
merge_pull_requestdelete_file |
Opening a PR is deliberately ungated. Branch, file and PR creation change nothing
that is running — they produce the document a human reviews. Stopping to ask "may I open
a PR?" leaves the incident burning while waiting for permission to do the one thing that
helps. Merging is what reaches production, and a human does that in GitHub.
Only 11 of 44 tools are enabled: re-sending all 44 schemas every turn is
what pushed investigations into rate-limit failures. |
| raw-file | 1 tool | none | A pure read at an exact git ref. Gating it would recreate the deadlock it was built to break. |
| notion | all | API-delete-a-block |
@destructive classified API-post-search as a mutation by its
REST-shaped name and stalled every postmortem before it could find the database. Writing
a postmortem is this agent's job; only deletion is gated. |
| n8n-builder | all 7 doc tools, +18 with an API key |
11 named tools | n8n_create_workflow is not gated. Its own description says "Created
inactive" — so creating one changes nothing that runs, and it is this agent's entire
purpose. There is no activate tool: activation happens inside
n8n_update_*, which is why both updates are gated even though "update"
sounds tamer than "activate". |
| n8n-tools | all live: harness_selftest, list_automations |
@all |
Every tool here reaches a real human or runs a real automation. The one server where "gate everything" is the correct policy rather than a shortcut. |
Why the gates are structural
A gate is not an instruction the model chooses to follow. It is declared in the agent manifest as
requireApprovalForTools, and the harness refuses to execute the call until a
human decides.
A paused turn and a finished turn both report status "done" — the pending gates are
in requiredActions. Treating that as the end of the session is the subtle bug that
would silently strand every approval, so the surface detects it specifically.
Said plainly: what the gates do not cover
"Never call alerting_manage_rules with update" and "never kubectl a fix
into place" are instructions in the prompt, not access controls. The harness gates by tool
name; it has no argument-aware authorization. A model error or a prompt injection could in
principle act outside those rules. That is a real limitation of the design, not a detail to leave
out of the diagram.
What runs where
Everything that touches a cluster or holds a vendor credential runs on the local host. The only thing deployed publicly is this page, which is static and can invoke nothing.
- This explainerstatic HTML · no build step · no env vars · no secrets
- TrueForge harnessnpx @truefoundry/trueforge · :8790 · runs both agents
- Orchestrator:8080 · alert admission, sessions, approval audit, Slack surfaces
- kind clusterdemo service · ArgoCD · kube-prometheus-stack
- MCP bridges:8100–8106 · stdio servers fronted over HTTP
- n8n + n8n-mcp:5678 · :8105 · the MCP tool hub
- Model proxy:8120 · absorbs the 429s the harness does not retry
- OpenAIgpt-5-6-terra, via the local retry proxy
- GitHubhosted MCP · revert PRs · Qodo App reviews the diffs
- Slacktwo apps, two Socket Mode connections — outbound, no public URL
- Notionpostmortem database
Two Slack bots, not one. SRE-Oncall owns the incident channel and hears about alerts; Automation-Agent lives in its own channel and routes to the automation agent. Each is a separate Slack app with its own identity and its own Socket Mode connection, and each is scoped to its own channel — a mention fires in any channel a bot is in, so without scoping two bots in one room would both answer.
The loops
Safety model
.env. This page holds none and needs none.What is real, and what is not
A demo that overstates itself is worse than a smaller one that does not. This is the honest split.
xcode-select needs. Rather than pretend, the runbooks are read out of git through the raw-file MCP server, and the routing tables live in the prompts. The skill definitions are correct and stay in the repo — on a host where the sandbox works, re-attaching them is a one-line change.