Chapter one runs a mission end to end in ninety seconds. The other seven open the rooms it passed through: the council that reads a plan, the code map, the crew, the rules, the routines, the money and the audit trail. Start anywhere, leave any time. No signup, no wall, nothing installed.
Guided simulation. Scripted data, real product flow.
Step 1 of 6
This is Carth in 90 seconds. Nothing here is live, everything here is how it works.
You are about to do the two things a person actually does in Carth: write the brief, and decide whether the work ships. The crew does everything in between, and you watch it happen from the same screen your team would.
Pick your enginesBrief a missionApprove what ships
Carth routes each role to a provider, on your own keys. Pick the two here and the rest of this walkthrough is labelled with them.
In the catalogue only
GeminiGrok
Their models are tracked and priced in your catalogue, and no agent is dispatched to them until an adapter lands. Provider names are the trademarks of their owners.
A role set to Codex falls back to Claude when the Codex command is not installed on that machine, and the fallback is recorded as an event. That is the real behaviour, not a shortcut taken for this page.
Step 3 of 6
Write the brief.
One plain sentence about the outcome you want. Change it if you like: it is carried through the rest of the walkthrough.
Work type FeatureRepository bakery-siteAgents at once 3
In your cockpit this dispatches a planner that reads the repository read only and comes back with tasks, assumptions and rough cost before a penny is spent. Here it plays a recording.
Step 4 of 6
The crew takes it from here.
Planning, building, reviewing. You can close the tab and come back to a decision waiting for you.
m_4f21c9
Build a landing page for our bakery with an order form
review diffs
Planning
Building
Reviewing
PlannerClaude
BuilderClaude
ReviewerClaude
Analystnot on this one
What the crew is doing
00:04Planner read the repository read only and split this into 3 tasks.
00:11Plan approved. 3 tasks dispatched, 3 agents working at once.
00:26Builder is working in an isolated worktree on mission/m_4f21c9/t1.
01:12Evidence saved to the mission: two drawings and a test run.
01:48Reviewer found one issue in the order form. Builder fixed it.
02:05Diffs are ready for you. Nothing has been committed or pushed.
In the real cockpit the plan is a gate too: you read it and approve it before any agent spends anything. This walkthrough approves it for you so the whole thing stays under two minutes.
Step 5 of 6
Your gate. Pictures first, the diff underneath.
The approver is not always the person who reads code, so the work shows itself before it argues.
m_4f21c9
Build a landing page for our bakery with an order form
review diffs
3 files changed+214-6mission/m_4f21c9/t1
Evidence
What the crew saved to show its work.
The landing page, drawn from your lineThe order formTests, all passing
Merge and deploy?
Recorded with your name and the time.
Builder adjusting the order form...
Round 2. You sent it back, the builder changed it, and the reviewer passed it again. That loop happens inside the harbour as often as it needs to, and you only see the result.
Approve commits the work on its own branch. Merging to develop and promoting to production are two more presses after this one, each held by a role and each recorded. Nothing here merges itself.
Step 6 of 6
Shipped.
This is what your real cockpit does, with your repos and your rules.
You wrote one line. The crew planned it, built it and reviewed itself before you were asked anything.
Every build happened in an isolated worktree. Your main branch was never touched.
Nothing moved until you pressed approve, and the record carries your name and the time.
That is chapter one. The rest of the tour opens the other rooms of the cockpit: the council that reads a plan before it runs, the map the crew reads, the crew itself, the rules it works inside, the work that happens on a clock, what all of it costs, and what is written down.
Chapter 2 of 8 3 beats
Before you approve a plan, put it to a council.
The plan you just approved had one model behind it. Council hands that same plan to three seats at once, each of them reading it alone, and gives you back where they agree, where they dissent and the reasoning behind each position. On Enterprise.
One plan, three seats
You pick the models that sit. Each seat receives the same mission plan, reviews it alone with read only tools, and is never shown another seat's answer, so nothing one of them says can pull the others towards it.
The struck links are the point. There is no channel between the seats, so the three reads meet for the first time on your screen.
What comes back
The three answers are put side by side by arithmetic and not by a fourth model. Where the verdicts match is the agreement, where they differ is the dissent, and every position stays attributed to the seat that held it.
Claude OpusConcernsTask 2 posts the order form to a route that does not exist yet, and nothing in the plan creates it.
Claude SonnetApproveThree small tasks, and the tests named in task 3 cover the form the brief asked for.
CodexConcernsThe same missing route, and a public form with no rate limit in front of it.
Two of the three raised the missing route, so it is shown as two seats raising it rather than as a score.
This one split, and the split is what you came for. One seat would have shipped the plan and two would not, and the thing worth reading is the route none of them was told about by the others. One agent agreeing with itself is not a second opinion.
Who sits, and how often
This install's council
SeatsThree, always
Sitting todayClaude Opus, Claude Sonnet, Codex
May hold a seatClaude and Codex
This month4 of 20 council runs used
The seats are yours to change. Only an engine that can actually run an agent may hold one, because a model that cannot read your repositories and report back is not reviewing anything. The three may not all be the same model: identical models produce correlated answers, which is an expensive illusion of consensus.
A council is minutes of paid work on three engines, so the allowance is counted per install, per calendar month. The count is read from the run log, which is also the record of what was asked and when.
Council is on Enterprise. It reviews the plan, before anything is built, so it is convened from a mission holding a plan at its gate; the reviewer you met in chapter one is the one that reads the finished work. The verdicts above are scripted, like every other number in this walkthrough. The three seats, the refusal and the monthly allowance are the product's.
Chapter 3 of 8 3 beats
The crew reads the map before it writes a line.
Carth builds a graph of your repositories and keeps it. Every mission starts from that map, so an agent is not paying to rediscover your codebase each time it is asked a question.
The map
One dot is a file. One colour is one community the graph found. The grey lines are the few places one part of the codebase reaches another.
62Backend endpoints
143Calls from screens
138Connected OK
5Need a look
One community
The graph groups files that change together. This one is the ordering flow: eighteen files, one checkout, and the two thin lines leaving it are the only places the rest of the codebase touches it.
Ordering18 files, 2 edges out
One node
POST /api/orders
CommunityOrdering
Called by24 screens
VerdictConnected OK
Most used endpoints are listed first, because changing one of them touches the most screens. Calls with no matching route are listed separately, and they are usually old code.
Code has three views in the cockpit: this graph, release drift between develop and main, and the raw exports behind both. The layout drawn here is sampled from a real map of 2105 nodes in 134 communities; the numbers beside it are scripted, like everything else in this walkthrough.
Chapter 4 of 8 3 beats
Meet the crew, then name one of your own.
A mission hands out generic roles: planner, implementer, reviewer, evaluator. The roster decides who fills them. Which engine answers each role is set one room away, in Settings, and the last beat here says where.
The roster
Roles are generic. The agents that fill them are not: each has a category, an expertise and a set of keywords, and a mission dispatches the one whose keywords match the task.
SherlockInvestigationReads before it touches. Reproduces the problem first.
MaestroImplementationBuilds the change inside its own worktree.
VeraEvaluationThe gate before your gate: passes or sends back.
CipherSecurityApplication security. Wins a review on a keyword match.
WardenSecurityInfrastructure and operations.
A mission deploys Cipher carrying Cipher's own system prompt, not a faceless reviewer.
Your own agent
Describe the specialist you wish you had. A read only agent drafts the persona, and nothing is saved until you have read every field.
You wroteAn agent that knows our checkout and refuses to let a price be computed in two places.
Name
Sable
Category
Review
Expertise
Checkout, pricing, order totals
Match keywords
checkoutpricetotalcart
Model
Category default Sonnet 5
Guardrail it carries
Never re-implement an order total. Use the canonical one.
Save agentDrafted read only. You edit every field before it exists.
The role matrix, which lives in Settings
Crew is the live flow and the roster, and nothing else. The table below is a panel of Settings, opened from the account menu at the foot of the sidebar. It is shown here because it is the other half of the same story: the roster says who the agents are, the matrix says which engine answers each role.
Mode: BalancedEco and Max quality fill the same table in one press
Which engine answers which role
Role
Engine
Model
Effort
Planner
Claude
Opus 5
high
Implementer
Codex
GPT-5.6 Terra
provider default
Reviewer
Claude
Sonnet 5
medium
Evaluator, Vera
Claude
Haiku 4.5
low
This is the same choice you made in chapter one, written out in full. A per mission override beats the table, and a role with no entry keeps its own default.
Crew has two views: the live flow of missions in flight, and the roster. The roster edits who runs what, so it is admin only, and a non admin who opens Crew lands on the live flow instead. Settings is admin only for the same reason, which is why the role matrix sits behind the account menu rather than on the sidebar.
Chapter 5 of 8 3 beats
The lines the crew works inside.
Three different things live here, and the difference between them is the whole point: a guard blocks an action, a rule is guidance every agent carries, and a skill is a command you can run.
A guard, and a skill
GuardNever touch main
PreToolUse / Bash
A script that runs before the tool does. Exit 2 blocks the action, exit 0 allows it.
On, enforcing
Skill/release-notes
Markdown command
Turns yesterday's merged work into release notes. Available in new console sessions the moment you save it.
On
Turning a guard off asks you to confirm, because it stops enforcing. Turning a skill off just renames the file, and turning it back on restores it exactly.
An operating rule
Ruleverify, do not guess
developmentships with the product
Read the code, run the command, check the config. Never present a plausible guess as a fact.
On, in every agent's context
Enabled rules are folded into every agent's prompt, on missions and in the console. A rule is guidance rather than enforcement: the guard above is the thing that actually stops a command.
A mission, blocked
Blocked by Never touch main
git push origin main
The builder tried to push. The guard stopped it before the command ran, the attempt is on the mission's trail with the time, and the work is still sitting in its worktree waiting for you. Nothing was pushed.
Parked by today's budget
$20.00 of $20.00 used
The queue reached the day's cap, so it parked the next mission with a budget event and told you, instead of launching it and finding out afterwards. A budget of zero means no cap, never "block everything".
Underneath all three there is a list agents cannot argue with. Write capable runs are refused git push, git commit, git reset, rm, sudo, docker, npm publish and the rest of the danger list, and they are confined to their own worktree. Every commit, merge and deploy is a separate human press.
Chapter 6 of 8 3 beats
Standing work, on a clock.
Turn one on, set the time, and check back that it ran. Three ship with the product and all three are off until somebody turns them on.
A routine
Security watchOn
Checks your dependency versions against current advisories and files anything high or medium to the backlog. Report only.
WeeklyMonday07:00
Next runMon 18 Aug, 07:00
Last runMon 11 Aug, 07:00
ResultRan ok
History6 runs kept
Every run keeps a small report, and a run that dispatched a mission links to it.
What it filed
Monday's run found two advisories and filed them. One was picked up as a mission, and that mission joined the queue exactly where a mission you wrote yourself would have joined it: waiting for a person.
m_7c04a1
Dependency advisories, week of 11 August
awaiting approval
Filed by Security watch, not by a person. It is waiting at the same gate.
What a routine cannot do
It cannot merge.
It cannot deploy.
It cannot promote to production.
It can look, report, and file work for a person to approve.
A routine skips its turn when the previous run is still going, or when the day's budget is already spent. If the machine was off at the scheduled time, it runs once on the next tick rather than catching up on a week at once.
The three that ship are Memory lint, daily and report only; Security watch, weekly and report only; and Server health, daily, which probes the endpoints you configured and uses no AI tokens at all. Beyond those, an admin can write a routine that runs an agent mission or a plain local command on a schedule.
Chapter 7 of 8 3 beats
What it cost, per mission, per model, per role.
Actual dollars where the run reported them, priced from your model catalogue where only tokens came back, and marked as an estimate when it is one. A number is never invented to fill the column.
The last seven days
$12.84 est.Spend$9.41 actual and $3.43 estimated
1.9MTokens1.4M in, 0.5M out, across 84 runs
$3.10Today's burn, of a $20 budget$16.90 remaining
~$21.40 est.SavingsAgainst running everything on Opus 5
The savings tile is an estimate and says so. It compares what you spent with what the same tokens would have cost on one baseline model you choose.
By mission
Every run in the project, gathered onto the mission that asked for it
Mission
Runs
Tokens in / out
Cost
Bakery landing page and order form m_4f21c9
9
412k / 98k
$2.41
Maestro, GPT-5.6 Terra codex
4
240k / 61k
$1.47
Planner, Opus 5 claude
2
121k / 24k
$0.88
Vera, Haiku 4.5 claude
3
51k / 13k
$0.06
Dependency advisories m_7c04a1
3
96k / 21k
$0.42
Rewrite the pricing page m_2b81e0
6
188k / 44k
subscription
That last row is real behaviour, not a gap in the data: a run on a Max subscription reports no dollar figure, so the cockpit writes "subscription" rather than a price it does not have.
By model, and whether it was worth it
Tokens turned into work that was kept
Model
Runs
Cost
Tokens / completed run
Waste
GPT-5.6 Terra
22
$5.90
34k
4.1%
Opus 5
12
$3.60
41k
0.0%
Sonnet 5
41
$3.10
19k
2.6%
Haiku 4.5 best value
9
$0.24
7k
0.0%
Waste is tokens spent on runs that failed or were cancelled. A run that failed, was retried and then succeeded is counted as a retry rather than waste, because the work did land. The same table exists per role.
Spending covers every run in the project, not just missions: console turns and the advisor are in the same figures. It is read only and computed when you open it. Where a model has no price in your catalogue the row says "no rate" rather than showing a zero.
Chapter 8 of 8 4 beats
Nothing crosses the opening without a name on it.
The last thing worth knowing is the part that does not change with the brief: what the crew cannot do, and what is written down the moment a person decides something.
11 Aug 14:06Maya merged to develop m_4f21c9branch mission/m_4f21c9/t1
11 Aug 16:31Leo promoted to production bakery-siterelease approver
Three presses, three names, three times. The audit log is durable: clearing the missions does not clear it.
Your keys, and what an agent is handed
Keys are write only. The cockpit will tell you a key exists. It will not hand the value back, not to you and not to an agent.
Agents are spawned without your secrets. No cockpit token, no cockpit home, and no variable whose name looks like a token, a secret, a key or a password.
Text is scrubbed before it is stored. Private keys, JWTs, API keys, connection strings and email addresses are stripped on the way into the log, not on the way out of it.
Where the work happens
One worktree per task. Each code task builds on its own mission branch in its own checkout, so parallel agents never collide and your main branch is not in the room.
An agent cannot be pointed at the cockpit. If a working directory would land inside the mission store or the secrets file, the process is refused before it starts.
One container per company. Your repositories, your keys and your mission history stay inside it.
The three gates, in order
The planYou read what it intends to do, and what it will cost, before a penny is spent.
The reviewYou read the pictures, the tests and the diff, and you approve or you send it back.
The shipMerging and promoting are two more presses, each held by a role, each recorded.
That is the whole tour.
Seven chapters, one argument: the crew moves fast inside the harbour, and a person with a name stands in the one opening.
You wrote one line and the crew planned it, built it and reviewed itself before you were asked anything.
It read the map first, worked under your rules, and could not reach past its own worktree.
Everything it cost is on one page, per mission, per model and per role.
Nothing moved until somebody pressed approve, and the record carries a name and a time.
Every plan starts with a short application, because every company we take on gets a setup call from a person. We reply within two business days.
All chapters
Escape opens and closes this list.
The same thing, in writing
The whole tour, in thirteen paragraphs.
Everything the guided version shows you, without the guide: the mission run first, then the seven rooms it passed through. The data is scripted; the behaviour is the product's.
Welcome
Carth is a cockpit where a crew of named AI agents plans, builds and reviews product work, and a person approves what ships. You do two things in it: write the brief, and decide.
Pick your engines
Each role runs on a provider you connect with your own key. Claude and Codex run missions today. Gemini and Grok sit in the model catalogue, tracked and priced, with no agent dispatched to them yet. A role set to Codex falls back to Claude when the Codex command is not installed, and the fallback is recorded.
Brief a mission
You write plain sentences about the outcome you want, pick a work type and the repositories in scope. The planner then reads those repositories read only and comes back with tasks, assumptions and rough cost, and you approve that plan before anything runs.
Watch the crew
Approved tasks fan out, three agents at a time by default. Each code task builds inside its own git worktree on a mission branch, so parallel work never collides and your main branch is untouched. A reviewer reads the result, runs the checks and attaches evidence, and weak work is sent back inside the harbour before it reaches you.
Your review gate
The gate opens with the pictures: the drawings and test runs the crew saved, then the diff summary, then the diff itself. You approve or you send it back. Sending it back adds a fix round to the brief and the builder works again.
Shipped
Approving commits the work on its branch. Merging it to develop and promoting it to production are separate presses, each held by a role and each recorded with a person, a moment and the evidence behind it.
Council
You pick the models that sit. A plan waiting at its gate can be put to a council of three seats: each one reviews that same plan alone, with read only tools, and is never shown another seat's answer. The three answers are then put side by side by arithmetic and not by a fourth model. Where the verdicts match is the agreement, where they differ is the dissent, and every position stays attributed to the seat that held it, with its own reasoning. One agent agreeing with itself is not a second opinion. The seats are configurable, only an engine that can run an agent may hold one, the three may not all be the same model, and the run allowance is counted per install per calendar month. Council is on Enterprise.
Code
Carth keeps a graph of your repositories: files and symbols as nodes, gathered into the communities that change together, with the endpoints your screens call joined to the routes that answer them. Missions start from that map instead of rediscovering the codebase. The Code destination also shows release drift between develop and main, and the raw exports both views are built from.
Crew
The roster is who the agents are: a name, a category, an expertise and keywords, so a task about security dispatches Cipher rather than a faceless reviewer. You can describe a specialist of your own, have the persona drafted read only, edit every field, and save it. Which provider and model answers each generic role is the role matrix, with Eco, Balanced and Max quality presets that fill the whole table in one press, and it is a panel of Settings rather than a view of Crew: Settings and its sections are reached from the account menu at the foot of the sidebar.
Rules
Three things that are often confused. A guard hook is a script that runs before a tool does and can block the action outright. An operating rule is guidance folded into every agent's prompt, on missions and in the console. A skill is a markdown command you can run. Turning a guard off asks for confirmation, because it stops enforcing. Underneath all three, write capable runs are refused a fixed danger list including git push, git commit, rm and sudo, and are confined to their own worktree.
Automations
Routines are agents and commands that run on a schedule: turn one on, set daily, weekly or monthly with a time, run it now, and read each run's report. Three ship with the product and all are off by default: memory lint, security watch and server health. A routine skips when the previous run is still going or the day's budget is spent. It can look, report and file work for a person to approve. It cannot merge, deploy or promote.
Spending
Tokens and price across every run in the project, missions, console turns and the advisor alike, by range, mission, model, provider and role. Actual dollars where a run reported them, priced from your catalogue and marked as an estimate where only tokens exist, "no rate" where a model has no price, and "subscription" where the run was on a subscription that reports no dollar figure. Efficiency joins usage to outcomes: tokens per completed run, cost per accepted mission, and waste, which counts failed and cancelled runs but not a retry that then succeeded.
Security
Provider keys are write only: the cockpit confirms a key exists and never returns the value. Agents are spawned without the cockpit's token, home directory or anything shaped like a secret, and cannot be pointed at the mission store or the secrets file. Text is scrubbed of keys, tokens, connection strings and email addresses before it is written to any log. Each code task builds in its own worktree, each company runs in its own container, and the audit log of who approved what and when is durable: clearing the missions does not clear it.