

The agent gets the repository instead of a prompt: the conventions file, the expert subagents that own each domain, the tools it may call, and a plan specific enough that a wrong assumption shows on paper. What it reports gets checked: the build, the tests, the applied delta, the digest on the box, a claim is never the evidence. What runs unattended and what waits for a person is set per environment before the first run.

Harness, Skills, Hooks
Every session writes back what it learned: a correction becomes a rule, a procedure becomes a skill with its failure modes attached, an incident becomes a line in the lessons log beside the code. Guard hooks enforce the limits. What belongs to the codebase stays in the repository, what is general moves into the library other projects read from. You get every application in source code, backups from day one, and a log of every conversation.

Worktime, Secrets, Dashboard
The result runs in daily operation: Worktime runs the time ledger and billing over its own MCP server, Secrets holds configuration values and keys that deployments read from, and the customer dashboard runs as the portal on the Hetzner box. This practice’s Azure and Hetzner skills run the CI/CD, from the GitHub Actions pipeline through to the blue/green deploy.
Full conversation log
every conversation with the agent is kept alongside the code it produced
Your sign-off decides
a draft waits for your review before anything saves to a live system
Invoice traces back
every figure on the invoice traces back to a real, recorded work session
Site confirms live
the public site itself confirms an update before it counts as live
From Plan to Reliable Operations
Stocktake & Plan 3 steps
01
Understand First, Then Let Go
Before the AI works on its own, everything worth knowing about the business gets written down first: what's running today, where, which databases exist, and how logins work. An open question gets marked as open rather than guessed, and a password gets described only by where it's kept, never written out. What's left behind is a document the next session picks up from.
02
Work Split Into Roles
The work gets split up like a small team: one role writes, one checks, one translates, one researches, one publishes, each with its own brief. A review decides whether something needs a second pass, capped at two, so no loop runs forever. The AI's full attention goes where real judgement is actually needed.
03
Plan First, Then Act
Every task gets the same clear shape before it starts, so a wrong assumption shows up on paper instead of in the build. That's how the 83 lessons in the business class are written, for one. The same discipline applies to a support request: the code gets looked at first, so the request names the real file instead of a guess.
Knowledge Built In 3 steps
04
The Safety Net Is Part of the Work
The safety net the training teaches is itself in use: specialists for each area, a team that checks every change automatically from three angles, and fixed house rules the AI has to follow. All of it is used live in the lessons. Every instruction lives in exactly one place, and the work runs reliably because of it.
05
Every Mistake Becomes a Rule
Every mistake made once becomes a rule the AI can't break again. One example: when domain records get updated, a forgotten entry would otherwise get silently deleted, so a backup copy gets made automatically first. A recurring task can go to the business side as a simple button in Microsoft 365, no technical tool required.
06
Rules That Enforce Themselves
House rules get enforced automatically instead of just written down and hoped for. A check at the end makes sure nothing slips through. A long-running version can measure the website's load speed on its own, on a schedule, with tight limits on what it's allowed to touch, so nothing risky happens unattended.
In Daily Use Here 2 steps
07
Two of Our Own Tools in Daily Use
Two of this practice's own tools actually run here every day, not just as an example. One tracks working hours and turns them into billing. The other holds every password and setting in one safe place that every other system reads from automatically, instead of someone copying it by hand.
08
Every Update Proves Itself
Before an update counts as live, it has to prove itself: what's now running gets compared against what was meant to go live, and the last check is whether the website actually answers. That matters, because once everything looked healthy while the site was still unreachable for 25 hours, until someone checked properly.
Visible and Billed 2 steps
09
The Whole Operation at a Glance
A dashboard shows the whole operation at a glance: which app is running, whether it's reachable, how it's configured, and what the work log records. It runs in daily use itself, in the same look as the public website, not just as a sample.
10
Test Honestly, Bill Honestly
What actually got tested gets shown honestly, and what didn't stands out as a gap instead of pretending everything was covered. Billing is built from the recorded working hours, so every figure on the invoice traces back to a real session.
Frequently Asked Questions
Does this only produce application code?
The same harness produces course and class material for any subject, through a factory that composes slides with overflow checks, generates narration per slide, composites annotated photos and renders avatar video. It produces multimedia through a media service exposed twice, over HTTP for what a person clicks and over MCP for what a model calls. It also produces operations output, including billing rebuilt from the recorded time ledger.
How is agent behaviour bounded in a real repository?
By hooks and gates. One guard hook caps the instruction file at 90 lines so it stays short enough to be read on every turn, another prints the skill roster on every prompt, and a verification gate sits behind them. Where a tool is irreversible, an approval gate goes in front of it: the governed MCP exercise hands over a live server with six tools including one destructor, and requires boundary rules plus a stress test against adversarial prompts.
Can a model write directly to a live business system?
A draft tool returns an interactive form rendered inside the chat client, prefilled with what the model proposed, while the model itself has no write tool and saving runs through a separate path triggered by the person looking at the form. Read-only integrations stay read-only, and a skill names the git-ignored file that holds a value instead of the value.
How does a deploy prove that anything actually shipped?
It compares the digest just pushed against the image on the box, compares the running container image id against that image, and finishes with a check that the public hostname responds, reporting an unchanged digest as a deliberate no-op. The same checks are available outside a deploy through the health and identity endpoints, so the question can be asked at any time.
