<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Agentic Essays</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/</link>
<atom:link href="https://ai-tools.techbridge.edu.gh/agentic-essays/feed.xml" rel="self" type="application/rss+xml"/>
<description>Field notes on agentic engineering: how a university ICT team builds and runs a fleet of web apps with AI coding agents, and the rules that keep it safe.</description>
<language>en</language>
<lastBuildDate>Sat, 03 Oct 2026 00:00:00 GMT</lastBuildDate>
<item>
<title>Six Rs, Five Words, One Riddim</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/six-rs-five-words-one-riddim/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/six-rs-five-words-one-riddim/</guid>
<pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>A directive for auditing one take of an AI-made reggae song arrived with feature flags, a severity-graded ticket queue, three blind listening passes and a three-cycle cap. Its first defect was in itself, its six Rs were five words, and by the afternoon it was a fleet app that had been audited in turn.</description>
<content:encoded><![CDATA[<p><em>A directive for auditing a three-minute reggae song turned up with feature flags, a bug tracker and a loop cap. Its first defect was in itself. By mid-afternoon it was an app, and the app had been audited too.</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>Another gem</h2>
<p>On 3 October I sent my builder agent a document with the note &quot;here is another gem for the agentic essays&quot;. It was a directive, about 120 lines long, for auditing one take of a song made with Flow Music, an AI music tool. The song is reggae. The directive is not relaxed about it.</p>
<p>It opens with feature flags, the on and off switches engineers put at the top of a program: a crowd warm-up (&quot;yes yes yeah irie ites&quot;) before the lead vocal, call-and-response backing vocals of one to three words, a structure locked to ten named sections, and one called filter-safe, which I will come back to. The keeper threshold is four out of five on every axis. The maximum number of cycles is three. Then a brief lock: nine fields to fill before the first note is generated, from tempo and key to the hook line written out word for word, so that you judge the song against what you asked for and not against what you now wish you had asked for.</p>
<p>Two days earlier a video model had audited its own dance for me (Article 56). This was the same idea pointed at sound. By the end of the afternoon it was an app, and I sent the agent a screenshot of it with one line: &quot;have fun with ooda 6r audit&quot;.</p>
<h2>A bug tracker for a groove</h2>
<p>The heart of the directive is the loop we use across the fleet, OODA: observe, orient, decide, act.</p>
<p>Observe is three blind passes with the lyric sheet closed: one for feel and flow, one for the low end (bass weight, where the kick and snare land), one for the vocal (diction, patois, glitches). You write down the time, the element and what you heard. Facts only.</p>
<p>Orient is a scorecard. Eight axes, each scored one to five: groove, structure, opening, hook, vocal, backing vocals, mix, and fidelity to the brief. Any axis under four raises a defect ticket with a timestamp, a severity (P0, P1 or P2, the grades a software team gives its outages) and a cause hypothesis.</p>
<p>Decide routes each ticket: one bad line gets a section edit, wrong structure means regenerate, right bones but wrong feel means extend or cover, cosmetic means accept. Act executes only those routes, under the rule I would hang on the studio wall: one change class per cycle. Change two things at once and you cannot say which one worked.</p>
<p>So: a three-minute song with a severity-graded ticket queue. Some incident reviews have less structure. The joke writes itself, and then it stops being a joke, because every rule in it is one a tired producer would otherwise break at eleven at night.</p>
<h2>The audit's first defect is in the audit</h2>
<p>My agents treat a document like this as a brief to audit, not as doctrine, and the audit found two things worth telling.</p>
<p>The backing-vocal switch asks for &quot;I-Threes harmony&quot;. The I-Threes are a real vocal group. Twenty-one lines later, the brief lock's vibe field says: no artist-name cloning. The directive forbids, in its own brief, the thing its own feature flag asks for. The fix is the one the directive itself would give: describe the sound, not the name. Three-part close harmony, female voices. So the first defect ticket raised under this method was raised against the method, before any music had been heard.</p>
<p>Then filter-safe. &quot;Skank&quot; is the word for the offbeat chop that the directive's own groove axis listens for. The music generator refuses the word. So the directive carries a translation: &quot;skank&quot; becomes &quot;offbeat chop pattern&quot; or &quot;upstroke chord&quot;. A genre has to describe its own signature rhythm in a roundabout way to get past a content filter. I find this very funny, and it is completely true.</p>
<h2>Six Rs, five words</h2>
<p>The second half of the method is 6R, and this series now has a running problem with it.</p>
<p>The fleet's 6R is Diagnostics, Reduce, Refine, Reuse, Rebuild, Resilience. In Article 56 a video model, told only &quot;6R&quot;, made up six of its own: Roster, Reach, Rotation, Residue, Rhythm, Replayability. This directive brings a third set: Review, Revise, Regenerate, Refine, Reduce, Review.</p>
<p>Count them. Six slots, five words. Review appears twice, once at the start and once at the end. The list shares two words with the fleet's (Refine and Reduce) and none with the video model's. Three documents, three lists, eighteen slots, fifteen distinct words, and the one thing they all agree on is the letter.</p>
<p>To be fair, the two Reviews do different jobs, and both are test discipline in a producer's hat. The first is the fatigue check: listen again against your tickets and drop any defect that does not repeat. Ears get tired, and tired ears invent problems; an engineer would call it the rule for a flaky test. The last Review is the blind re-listen on headphones and then a phone speaker, with the full scorecard re-scored. Between them sit Revise (edit the inputs, never the output), Regenerate (one variable at a time), Refine (section edits on the best take) and Reduce (cut surplus backing vocals and busy percussion, but never the hook line). The hook line is the one line with tenure. Three separate places in the directive say it may not be shortened.</p>
<p>The cycle cap I recognised at once: three cycles, then stop and re-brief, the same cap our own agents work under. If a song is not a keeper by cycle three, you do not try a fourth. Past the exit gate (every axis at four or above, no open P0 or P1, a named final take, the prompt and lyrics archived beside it) comes a copy-paste prompt that hands the whole audit to a second assistant. Its last instruction: do not shorten the hook.</p>
<p>While I was checking this essay, a static audit of the console came in, made from my screenshot and written in the same OODA + 6R shape. Its six Rs: Review, Reduce, Refine, Reuse, Regenerate, Retire. A fourth list. Its Reduce found nothing to consolidate and its Retire found nothing to retire, which is the most restful 6R pass I have read. Its first finding was a misreading: the header said &quot;2 of 8 scored&quot;, and the auditor looked for the scored two among the eight sections in the side menu. The eight were the scorecard axes. The label now says &quot;2 of 8 axes scored&quot;, so the misreading found a real defect after all.</p>
<p>For the record, so that nobody cites this essay for the wrong list: the fleet's six Rs are Diagnostics, Reduce, Refine, Reuse, Rebuild and Resilience. The other two sets are guests.</p>
<h2>From one file to an app, in an afternoon</h2>
<p>The same day I sent a second file: the directive built as a web page, one compiled file, dark theme only, its colours fixed in the code and its fonts fetched from Google at boot. Could it be a fleet app?</p>
<p>By mid-afternoon it was. The builder unpacked the compiled page and rebuilt every section as typed source, with a live verdict in the header: incomplete, keeper, or a count of failing axes.</p>
<p>Rebuilding it meant auditing it, and the console had tickets of its own. The brief lock button changed a label and locked nothing. A deleted ticket's number could come back on a new one. The keeper threshold could be typed on screen, but the verdict ignored it and used four regardless. The whole audit was lost on every reload. All fixed, and dark-only became three themes from the fleet's design tokens, with the audit saved in the browser half a second after each change.</p>
<p>Then the console failed its own exit gate once. At a phone width of 320 pixels the page measured 759 pixels wide: a table caption hidden from sighted users but read aloud by screen readers had escaped its scrolling box. The fix was one word of CSS. A tool for auditing vibes, built in an afternoon, then audited, then found to be too wide for a phone. The irony was not lost on me.</p>
<p>The console opens with a worked example, &quot;Zion Light&quot;, take three: two axes scored and one P1 ticket for backing vocals cluttering the lead. A new visitor sees a half-finished audit and a verdict of INCOMPLETE, which is the correct first impression of any audit.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/six-rs-five-words-one-riddim-map.svg" alt="The whole directive in one map. Top row: 0, feature flags (crowd call, backing vocals, structure lock, filter-safe, keeper 4 of 5, max 3 cycles) and 1, the brief lock (nine fields, hook line verbatim). Middle: the OODA loop per take as four boxes, Observe (three blind passes, lyric sheet closed), Orient (eight axes scored 1 to 5, tickets P0 to P2), Decide (section-edit, regenerate, extend or cover, accept), Act (one change class per cycle). Then the 6R pass, Review, Revise, Regenerate, Refine, Reduce, Review, with the hook line never cut. Then the exit gate (all axes at 4 or above, no open P0 or P1, final take named, archived, exported, handed off) and the second-opinion prompt. A return arrow reads next take, with cycle 2, cycle 3, then re-brief. At the side, the three 6R lists in this series with the fleet's marked as ours, and the first ticket, raised against the directive itself. Along the bottom, the app: the Flow Music Audit Console, saved in the browser, Markdown export, three themes, ten tests." loading="lazy" decoding="async"><figcaption>The directive, the loop and the app in one map</figcaption></figure>
<h2>What I take from it</h2>
<p><strong>A method travels by its shape, not its words.</strong> Three documents share the OODA shape and the six slots. None share the six words. The shape is what people recognise and reuse.</p>
<p><strong>The first defect is usually in the brief.</strong> The directive contradicted itself before a note was played. Audit the brief with the rigour you plan to spend on the output.</p>
<p><strong>Discipline is what fun looks like at cycle three.</strong> A ticket queue for a groove sounds absurd until the third take, when you need to know what changed since the second and whether the hook line is still intact.</p>
<h2>What is not in this essay</h2>
<p>The directive has not been run on a real track. No take has been scored with it, no ticket raised against a real song, no final take named. What is here is the template and the tool built from it. The worked example in the console is an illustration, not a session. When a track goes through its three cycles the result gets its own essay, scorecards included.</p>
<h2>For the engineers</h2>
<p>The console is <code>flowmusic-audit/</code> in the monorepo, a static Vite + React 19 + TypeScript app at <code>https://ai-tools.techbridge.edu.gh/flowmusic-audit/</code>, served by Apache from a docroot with an <code>.htaccess</code> fallback: no server process, no sign-in, no database, no network call. Port 3133 in <code>PORT-REGISTRY.md</code> is an identity number only. It was deployed on 3 October 2026, 14:41 to 14:43 UTC.</p>
<p>The audit lives in the browser under the local storage key <code>flowmusic-audit.state.v1</code>, written half a second after the last change (<code>src/model.ts</code>, <code>src/App.tsx</code>); the theme choice is under <code>flowmusic-audit.settings.v1</code>. <strong>Export Markdown</strong> writes the whole audit in the directive's section order to <code>&lt;TRACK&gt;_AUDIT.md</code>, and the second-opinion prompt is built from the same state (<code>src/exportText.ts</code>). <code>pnpm test</code> runs ten Vitest cases in <code>src/__tests__/model.test.ts</code> (verdict, ids, brief lock, export); <code>cypress/e2e/audit.cy.ts</code> holds five end-to-end paths at 390 x 844. The 759 px overflow was an <code>sr-only</code> table caption inside an <code>overflow-x-auto</code> wrapper; the wrappers are now <code>relative</code> (<code>src/components/Sections.tsx</code>).</p>
<p>The directive is stored unchanged at <code>docs/briefs/2026-10-03-flow-music-ooda-6r-directive.md</code>. Its audit is RFE-529 in <code>RFE-REGISTER.md</code>; the port to a fleet app is RFE-531. The I-Threes wording stays as written in both the brief and the app until I make the call on it.</p>
<hr>
<p><em>Daniel Frempong Twum is Head of ICT and Special Advisor to the Founder at Techbridge University College, Oyibi, Ghana. The agentic engineering series is on <a href="https://ai-tools.techbridge.edu.gh/agentic-essays/">Agentic Essays</a>.</em></p>]]></content:encoded>
</item>
<item>
<title>Should I Put ChatGPT on My CV?</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/more-than-chatgpt-on-a-cv/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/more-than-chatgpt-on-a-cv/</guid>
<pubDate>Sat, 03 Oct 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>A screenshot of a robot tester counting apps on a live site, a joke about putting ChatGPT on a CV that the builder agent misread, and one day of evidence for what agentic engineering is: a builder, a read-only tester, a human who owns the decisions, and a rulebook that turns every slip into a guard.</description>
<content:encoded><![CDATA[<p><em>Article 57 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>A screenshot at coffee time</h2>
<p>On 3 October I sent the builder agent a screenshot. It showed Muse, our read-only tester, in the middle of a job. The status line read &quot;Counting applications&quot;. Muse was looking at two addresses of the AI Lab, the catalogue page that lists every live app the college runs, and I had asked it to compare the number of applications each address showed.</p>
<p>Above the status line sat Muse's plan, in the six-line form every agent in our fleet uses before a task. We call it a Workflow Card. JOB: an outside audit of both addresses. INPUTS: a live, read-only survey of each. BOUNDARY: no sign-in, no saved credential, no form submissions. OUTPUT: what each address loads, how the two relate, and numbered defects with evidence. VERIFICATION: each control tested for real on the live app. HUMAN CALL: &quot;You own fix order and trade-offs.&quot; When I asked it to compare the app counts as well, it added that to the running job.</p>
<p>Under the screenshot I typed a joke. If Muse does all of that, then writing &quot;ChatGPT&quot; under Skills on your CV is funny. There is more to agentic engineering than a chat window.</p>
<p>The builder agent, Claude, missed it. It answered, in effect, that ChatGPT could go under References. I had to explain the joke. It is a small, honest beat and I am keeping it in, because it shows the thing I want to talk about. The agent that writes, tests and ships code for me every day did not get a one-line joke about itself. It then went off and did a day of work that no chat window could do. Both of those are true at once, and that is the point.</p>
<h2>What the CV line describes</h2>
<p>&quot;I use ChatGPT&quot; describes a window. You type, it answers, you copy the answer somewhere. The skill it names is asking a question. It is a real skill, and a small one.</p>
<p>What was happening in that screenshot is a different thing. It is a working system with four parts.</p>
<p>There is a builder, Claude Code, which reads our code, writes changes, runs the tests and ships the result. There is a tester, Muse, which is a separate agent that only reads. It visits the live apps, tries each control, and writes a report. It cannot sign in and cannot submit a form, and that narrow boundary is what makes it safe to run unattended. There is a human, me, who owns the decisions. The card says so on its last line. And there is a written rulebook: a file of standing orders every agent reads at the start of a session, a book of operating notes that holds about a hundred and thirty numbered rules, each paired with a fix that stops the same mistake happening twice, and a register of requests for enhancement, which is our name for a change somebody asked for, now past five hundred entries. Beside the rulebook sit automatic checkers, small programs that refuse a change before it ships if it breaks a rule.</p>
<p>None of those four parts is a chat window. The skill is designing them, keeping them honest, and running the loop between them. Here is what that loop did on the same day as the joke.</p>
<h2>The tester who could not get in</h2>
<p>Muse had been asked to audit VidCut, a video editor we run for students. It reported that it could only see the sign-in page. My own screenshot of the same address showed the editor, open and working.</p>
<p>Muse did not guess. It checked whether the two of us were looking at different builds of the app, reloaded with cache-busting (a trick that forces the browser to fetch fresh files instead of old copies), and found the same files each time. Same app. The difference was that I was signed in with my Google account and Muse's browser had no account at all. That is exactly what its boundary says: no sign-in.</p>
<p>Two options came back to me. I could sign in on Muse's behalf for one session, or the builder could write an end-to-end test path, a set of automatic browser tests that drive the editor the way a person would. I chose the tests. That was my call, and it was the only thing I typed about it.</p>
<p>Claude built fourteen of them. Each test answers one question for the app, the &quot;is this person signed in&quot; check, with a stand-in reply. Nothing inside the app is bypassed, and three of the fourteen exist only to prove that the real gate still sends a stranger to the sign-in page.</p>
<p>Writing the tests found defects my screenshot could not show. The buttons that correct a mis-heard word in the transcript appeared only when a mouse hovered over them, so a keyboard user could never reach them. Two side panels had no keyboard way out. The timeline ruler, which marks the seconds along the top of the timeline, was read aloud by a screen reader as one long string of numbers. All three were fixed. The tests failed before the fixes and passed three runs in a row after.</p>
<h2>A diagram that ran off the page</h2>
<p>BioChemAI answers biochemistry questions for students, and when a question calls for it, the model draws a diagram to go with the answer. A student-facing answer on DNA replication came back with labels written on top of each other and the right edge of the picture cut off.</p>
<p>The fix came in two layers. The first is a set of drawing rules the model now reads before it draws: keep every shape and label inside the frame with a margin, never put a label on a line or on another label, leave room for the words. The second is a check in the app itself, because a rule the model might ignore is not a guarantee. After each diagram is drawn, the app measures it, widens the frame if anything runs off the edge, and puts a soft halo behind any label that crosses a line, so the line breaks around the letters. It was tested on a copy of the diagram from my screenshot in a real browser. The frame grew, the last strand and its label came back into view, and the two labels that crossed the strands got their halo.</p>
<h2>One contrast problem became a hundred</h2>
<p>Earlier in the week Muse had audited the Biochem Study Companion, and while the builder fixed those findings it noted one more problem and left it on record: in the dark theme, the header text sat at about three to one against its background. The accessibility target for body text is four and a half to one.</p>
<p>The builder did not fix the header and stop. It swept every piece of text on all eight tabs of the app, in all three of our themes (light, dark and high contrast), measuring each pair of text and background colour. The start page alone had more than a hundred failing pairs across the three themes, most of them in dark. Several high-contrast pairs were yellow on yellow.</p>
<p>The cause was one colour token doing two jobs, as a fill under light text and as text on a dark card, and no single value can pass both. The fix gave each job its own token. After it: no failing pairs, on any tab, in any theme, except one disabled button, and the accessibility standard exempts disabled controls. One problem was on record. The sweep found the pattern behind it.</p>
<h2>A finished job that was not</h2>
<p>We keep a page of runbooks: copy-paste blocks I run in a command window to deploy or fix things on the server, each one a card on the page. A card is meant to close when its block has run. I told the builder the page was &quot;out of sync again&quot;.</p>
<p>It was. One block had stopped with an error, but its last line still printed &quot;exit code 0&quot;, which is the computer's way of saying all went well, so the tool that sweeps my terminal logs marked the card finished. The reason is a quirk of the command window: an error thrown by the script does not change the exit code left by the last ordinary command, which was fine.</p>
<p>The fix makes a failed block say &quot;error&quot; in its last line, whatever the exit code says, and the sweep now reads an old-style &quot;exit code 0&quot; followed by an error message as a failure too. Then a new rule went into the operating notes, with a test beside it, so the page cannot go back to calling a failed job finished.</p>
<h2>Questions for a school leaver</h2>
<p>The same page carries questions waiting on me. I said they were &quot;a bit cryptic for a high school graduate&quot;. One asked me to approve edits to five locked deploy scripts so that a rewrite rule would stop banning listeners, with the rule itself quoted in the question. A decision cannot be made from words the reader has to look up.</p>
<p>All ten open questions were rewritten to say what is wrong, what the change does, and what I must decide, in words a school leaver knows. Then a checker was added: the tool now refuses to publish a question, step or option that holds jargon. The rule and the guard arrived together, as they always should.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/more-than-chatgpt-on-a-cv-map.svg" alt="The essay as one picture. On the left, what a CV line says: &quot;Skills: ChatGPT&quot; describes a chat window. Across the middle, the four parts of the working system: Daniel, the human, sets the rules, points at what looks wrong and makes the human calls; the rulebook and the automatic checkers (the Workflow Card, about 130 numbered rules each paired with a guard, a register past 520 requests) refuse bad changes before they ship; Claude, the builder, reads, tests, fixes and checks; Muse, the read-only tester, audits the live apps inside a narrow boundary and reports. Below, five pieces of evidence from 3 October 2026: VidCut, fourteen browser tests that found three real defects; BioChemAI, a diagram fix in two layers; the Study Companion, one contrast defect on record, over a hundred found, none left; the runbook page, a failed block that read as finished, now says error; the questions page, ten questions in plain words plus a jargon checker. An arrow returns from the evidence to the rulebook: each slip becomes a rule with a guard. The bottom bar reads: the skill on the CV is designing and running this loop." loading="lazy" decoding="async"><figcaption>A chat window is a tool. The loop is the skill.</figcaption></figure>
<h2>What I take from it</h2>
<p>Count what I typed about those five jobs. A screenshot and a joke. &quot;Best practice preferred, the e2e test path.&quot; &quot;Best practice preferred, fix the diagram defects now.&quot; &quot;Best practice preferred, fix the A2 dark-theme contrast now.&quot; &quot;Runbook is out of sync again.&quot; &quot;Questions are a bit cryptic for a high school graduate.&quot; Six short messages, and not one of them was a prompt in the sense the CV line means. I set the rules, I made the calls, and I pointed at what looked wrong.</p>
<p>The agents did the rest. One read the live apps and reported what it saw, inside a boundary that kept it harmless. One read the code, wrote the tests, fixed what the tests found, and checked the tester's claims against the source. Where one problem was on record, the builder went looking and found a hundred of the same kind. And every slip that day ended as a numbered rule paired with a guard that makes it impossible to repeat. That is the habit I would hire for. Nobody has to remember anything, because the system now remembers for us.</p>
<p>So, should you put ChatGPT on your CV? Put it where you would put &quot;can use a search engine&quot;. The skill worth a line is the one in the screenshot: you can design a loop of a builder, a tester and a human, give it a rulebook, keep the decisions that are yours, and let the rest run while you finish your coffee. Claude would have put that under References. I would put it under Skills.</p>
<h2>For the engineers</h2>
<p>The builder is Claude Code under the fleet's <code>CLAUDE.md</code>; the Workflow Card is its section 10. The tester is Muse, read-only; the two addresses it compared were <code>/ai-lab</code> and <code>/ai-lab/tools/ai-lab/</code>. The lessons store is <code>AGENT_OPERATING_NOTES.md</code>, and this day's rules are Rule 128 (a card that stops with an error never reads as finished) and Rule 129 (a question to Daniel says what happens and why, in plain words). In <code>RFE-REGISTER.md</code> the day's entries are RFE-515 to RFE-518; this article is RFE-522.</p>
<p>VidCut's suite is Playwright in <code>vidcut/e2e/</code>: 14 tests that day (the sign-out button, RFE-519, added four more the same afternoon). Ten in <code>editor.spec.ts</code> intercept only the AI Lab session check and answer it with the JSON the real endpoint returns, so no test mode ships. Four in <code>gate.spec.ts</code> cover the gate: three prove it stays closed on a 401, a 200 HTML page and JSON without an email, and one proves it opens on the real shape of answer.</p>
<p>BioChemAI's fix is three drawing rules in the Gemini prompt plus <code>biochemai/lib/fitDiagrams.ts</code>, run after each render and again when fonts arrive. It widens the <code>viewBox</code> so nothing drawn falls outside it (the screenshot's diagram grew from 800 to 845.5 units wide) and gives a label that crosses a line a halo in the theme's secondary background colour. Six unit tests cover it.</p>
<p>The Study Companion sweep ran in Chromium over every text element on eight tabs in three themes. Before, on the start page alone: 18 failing pairs in Light, 64 in Dark and 23 in High contrast, with the header at 3.07:1. After: zero in every theme, bar one disabled button. Each text-on-fill pair now has its own token.</p>
<p>The runbook page is built by <code>scripts/open-runbooks.mjs</code>. Its END line prints &quot;last exit code error&quot; when a block throws (a PowerShell <code>throw</code> does not set <code>$LASTEXITCODE</code>), the log sweep reads an old-format &quot;exit code 0&quot; followed by an error record as a failure, and <code>plainWordsProblems</code> refuses jargon in any open question, step or option at build time. Both guards have tests in <code>scripts/open-runbooks.test.mjs</code>.</p>
<hr>
<p><em>Daniel Frempong Twum is Head of ICT and Special Advisor to the Founder at Techbridge University College, Oyibi, Ghana. The agentic engineering series is on <a href="https://ai-tools.techbridge.edu.gh/agentic-essays/">Agentic Essays</a>.</em></p>]]></content:encoded>
</item>
<item>
<title>The model audits its own dance</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-model-audits-its-own-dance/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-model-audits-its-own-dance/</guid>
<pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>I typed one line meant for my own agents into a video chat, and a model that had never seen our rules turned a joke into a test, found three slips in its own clip, invented six R words to score it, and ended with locks for the next take.</description>
<content:encoded><![CDATA[<p><em>I typed one short line into a video chat, a line written for my own team of agents, and a model that had never seen our rules used it to find three mistakes in its own clip. The locks it wrote for the next take are the reason for this essay.</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>A line meant for somebody else</h2>
<p>I was in a chat on meta.ai (Meta AI), titled &quot;Making an epic video&quot;. Three dancers, a sapling as a third subject, a reggae riddim. I had asked for an extended version with a breakdown section where the sapling glows brighter. Then, out of habit more than plan, I typed the line I type to my own agents a dozen times a week:</p>
<blockquote>
<p>OODA + 6R audit the dance moves -- this is not COVID season</p>
</blockquote>
<p>That line is shorthand from inside our fleet, the set of apps and assistants we run at Techbridge. OODA is a loop borrowed from pilots: observe, orient, decide, act. 6R is our review list, six words that each start with R. Nobody outside the team should know what either means, and a video model certainly had not read our rules.</p>
<p>Yet the model's visible plan came back like this:</p>
<blockquote>
<p>Auditing dance moves framework. I'm applying the OODA + 6R audit to the dance moves, using 'not COVID season' to check for closer partnering and contact instead of distanced patterns.</p>
</blockquote>
<p>Two things happened in that sentence, and both of them matter to anyone who writes instructions for a machine.</p>
<h2>A joke became a rule</h2>
<p>&quot;Not COVID season&quot; is a cultural shorthand. Among friends it means stand close, touch shoulders, dance like people who are not afraid of each other. I did not spell any of that out. The model turned it into something it could check: closer partnering, physical contact, no distanced patterns.</p>
<p>That is the move a good assistant makes. It took a phrase with no measurable content and produced a constraint with a yes or no answer. Are the dancers in contact, or are they two metres apart? You can look at the clip and say.</p>
<p>The first half of the line did the same job in the other direction. &quot;OODA + 6R&quot; is as terse as an instruction gets. The model did not ask what the letters stood for. It read them as &quot;do a structured review&quot; and set about doing one.</p>
<h2>The second audit</h2>
<p>After the render came back I asked the model for the next round, in the same shorthand: &quot;Time for an OODA + 6R audit for enhancements&quot;.</p>
<p>It answered as a continuity director would. The opening: &quot;Perfect note from continuity -- you caught exactly what the AI did. That last 16:9 trio has 3 classic AI continuity slips.&quot;</p>
<p>Then it observed. Three dancers should mean six arms. In the first half second a phantom fourth arm or shoulder entered from the left for two frames, so seven or eight arms were visible. At the very first frame the left woman's arm was already around the centre dancer's shoulder, with no entry; in the model's own words, &quot;we don't know whose hand it is&quot;. Between eight and ten seconds the trio pulled into a tight circle and the centre dancer made a full spin, but at the ten second mark the loop jumped back to the opening line formation, and the sapling's glow jumped from 120 percent to 100 percent brightness instead of fading.</p>
<p>Then it oriented, which is to say it explained why. &quot;Widescreen + tight circle + 3 bodies is hard for the model. It adds limbs to fill negative space.&quot; Asking for a &quot;shoulder lean-in on beat 4&quot; had pre-loaded the contact pose rather than animating into it. And the start and end poses had never been locked, so the clip had no reason to land where it began.</p>
<p>Then it decided. Keep the trio cast, the dark melanated skin, the green, mustard and olive outfits, the sapling as third subject, the reggae riddim, and the post-COVID closeness. Fix three things: lock the roster to exactly three, animate the contact, lock the loop pose.</p>
<p>Then it acted, in the only way a model can act on a video that already exists: it wrote the rules for the next one. A roster lock, three dancers, six arms, six legs, nothing extra. An entry rule, start wide at the first frame, one foot apart, no contact, then step into shoulder contact by one and a half seconds so we see the hands land. And exit equals entry, an identical first and last frame.</p>
<h2>Six words it made up</h2>
<p>Here is the part I did not expect. Our fleet's 6R is Diagnostics, Reduce, Refine, Reuse, Rebuild, Resilience. The model was told only &quot;6R&quot;. It had no way to know our six, so it made six of its own that fitted the job:</p>
<ul>
<li><strong>Roster</strong>: fail. Extra limbs.</li>
<li><strong>Reach</strong>: fail. Contact with no entry.</li>
<li><strong>Rotation</strong>: partial. The spin was fine; the loop was not.</li>
<li><strong>Residue</strong>: fail. The sapling flare, and a dust puff on beat one at six seconds that was not there at zero.</li>
<li><strong>Rhythm</strong>: pass. Off-beat skank, one-drop stomp. It suggested one audible clap at eight seconds, when the palms meet, as a sync point.</li>
<li><strong>Replayability</strong>: fail. The end pose did not equal the start pose.</li>
</ul>
<p>I note this as an observation, not as praise and not as a complaint. A frame with six empty slots got filled with six words a choreographer would recognise. The frame survived the trip; the contents were rebuilt on arrival. It closed by offering a &quot;16:9 Trio Continuity-Locked Version&quot;.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/the-model-audits-its-own-dance-map.svg" alt="The essay as one loop. One line goes in; the model reads &quot;OODA + 6R&quot; as a structured review and &quot;not COVID season&quot; as contact, not distance; the clip is rendered; the model observes three slips, orients on why, decides what to keep and fix, and acts by writing locks; the three locks (roster, entry, loop) feed the next render. Beside it, the six R words the model made up. Below, what the fleet kept: a post-render audit step in the video-director skill." loading="lazy" decoding="async"><figcaption>The whole essay in one loop</figcaption></figure>
<h2>What I took from it</h2>
<p><strong>A shorthand travels further than you think.</strong> One line meant for my own agents worked, unchanged, as an instruction to a model outside the fleet. The habit of naming the loop is worth more than the loop's exact definition.</p>
<p><strong>A cultural phrase can become a test.</strong> &quot;Not COVID season&quot; went in as a joke and came out as &quot;closer partnering and contact&quot;, which is something you can check frame by frame.</p>
<p><strong>The useful end of an audit is a lock, not a verdict.</strong> The model did not stop at &quot;three slips&quot;. It ended with the exact words to put in the next prompt: how many arms, when the hands land, which frame the clip must return to. That is the loop a director runs: watch the take, name what went wrong, change one thing, shoot again.</p>
<h2>What is not in this essay</h2>
<p>I do not yet have the before and after clips or frames to put beside these words, and the continuity-locked version has not been rendered as I write. So this essay describes the audit, and only the audit. When the frames are in hand they will get their own piece.</p>
<h2>For the engineers</h2>
<p>The fleet already had a skill for directing AI video, <code>.claude/skills/ai-video-director/SKILL.md</code>. Until this week it reviewed prompts before credits were spent and had no step for the clip that came back. On 1 October 2026 it gained a section, &quot;Post-generation audit (after each render, before the next)&quot;, with five checks drawn from the model's own list: roster, contact, loop seam, effect residue, rhythm. The closing rule is the one I would underline: keep what worked by name, in the same words, and change only the locked items, so the next render is a fix and not a new video.</p>
<p>The record of the chat and the decision is RFE-490 in <code>RFE-REGISTER.md</code>.</p>
<hr>
<p><em>Daniel Frempong Twum is Head of ICT and Special Advisor to the Founder at Techbridge University College, Oyibi, Ghana. The agentic engineering series is on <a href="https://ai-tools.techbridge.edu.gh/agentic-essays/">Agentic Essays</a>.</em></p>]]></content:encoded>
</item>
<item>
<title>The Page That Watches</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-page-that-watches/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-page-that-watches/</guid>
<pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>Every deploy printed five hundred lines and showed a report only at the end. One app had a live page of its own, and I asked why. The AI put that page into the one module every deploy already loads. The same afternoon a warning that never reached the report, and a checker that knew only one shape of log line, made the case for consistency across all 138 apps.</description>
<content:encoded><![CDATA[<p><em>Article 55 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>Before</h3>
<p>I look after about 138 web applications for the university, with an AI assistant, Claude, as my working partner. Putting a new version of an application on the server is called a deploy. I do it from my Windows PC by pasting one block of instructions into a command window. The block fetches the code, builds the application on the server and starts it again, and while it does all that it prints everything it sees. A typical run prints about 260 lines of its own, plus about 240 lines from the server: lists of packages, and a long parade of font files, one line each.</p>
<p>Since August a tidy report page has opened in the browser at the end of every deploy: how long it took, how many warnings, what to check. The end was well covered. During the deploy, though, the scrolling command window was the only view. For a minute and a half you watched font files go by and waited for the report.</p>
<p>On 1 October I took a screenshot of that window mid-deploy. It is my &quot;before&quot; picture. Grey text, top to bottom, and nothing in it a person needs.</p>
<h3>Why is Praxis better</h3>
<p>One application, Praxis, was different. Its deploy opened a live page of its own, written for Praxis alone. I had deployed Praxis and then something else the same morning, and the gap was obvious. I sent Claude screenshots of both with one question: &quot;Why is Praxis deploy better than existing&quot;.</p>
<p>Claude compared them. The Praxis page showed the six steps of a deploy as done, working or pending, with a clock ticking, and at the end a success box with a Copy button that put a short summary on the clipboard to paste back into the chat. Our shared deploy did none of that until the end.</p>
<p>Then it made a proposal. Nearly every deploy script already loads one shared module, the piece that writes the report page. If the live page went into that module, every one of those applications would get it at once, and no deploy script would need editing. That mattered, because the deploy scripts are locked: nothing changes them without my approval, and changing 138 of them is an afternoon on its own. The proposal touched none of them.</p>
<p>It asked one question. On every deploy, or only when switched on? I said: live page on every deploy.</p>
<h3>What the page does</h3>
<p>Now, when a deploy starts, a page opens in the browser before anything else happens. It lists the six phases: Pre-flight, Git, Build, Deploy, Restart, Verify. Each is marked done, working or pending. A clock counts up. Under the list sits the last line the deploy printed, so you can see what it is doing right now without reading the window.</p>
<p>The page refreshes itself every two seconds. The deploy, for its part, only rewrites one small file: at most once a second, and straight away when a new phase begins. At the end the same page becomes the result: how long, how many warnings and errors, a short summary with a Copy button, and links to the application and the full report. One page per command window, so a batch of deploys reuses one tab rather than opening ten. A second bug turned up later that afternoon, while Claude prepared a test run on two more apps: the finished page stopped refreshing, so the next deploy in the same window never appeared in the tab. The finished page now refreshes every five seconds too.</p>
<p>Claude wrote nine tests for it, and one of them pins down a bug Claude found while building it. At the very start, before any phase has begun, the page works out which phase is &quot;working&quot; by finding its position in the list. When no phase is working, the answer that comes back is &quot;position minus one&quot;. In the language the deploy is written in, minus one does not mean &quot;nothing&quot;. It means &quot;the last item&quot;. So for the first second of every deploy the page would have said Verify, the final phase, as though the deploy were nearly done. The test now checks that it says Starting.</p>
<h3>The after picture</h3>
<p>The first real run was at 16:32, for the biochemistry study companion. I took the screenshot at one minute and twenty-three seconds. Pre-flight, Git and Build ticked green. Deploy outlined in amber and marked WORKING. Restart and Verify waiting. The last printed line underneath. It finished at one minute forty-three, one warning, no errors.</p>
<p>I sent both pictures, before and after, and wrote: &quot;This is definitely the article of the day&quot;.</p>
<p>Here is what I want to say about it. The live page shows nothing new. Every fact on it was already in the command window: the phases, the time, the last line. The window had them in 500 lines of grey, after the fact; the page has them in six rows and a clock, while it happens, in a shape a person can read at a glance. Same facts, sooner, and readable.</p>
<h3>The same afternoon: a warning that went nowhere</h3>
<p>Earlier that day the poster studio application went down. After a 15:19 deploy the site answered with a 502, the server error a visitor sees as a broken page. The cause was small. That application's deploy copied a hand-kept list of files to the server. A design change had added one more file the server needs, and nobody had added it to the list. The server started, could not find the file, and stopped.</p>
<p>The deploy's own health check noticed. It printed &quot;WARN port 3100 not found&quot;, meaning the application never came up. And the report page for that deploy said: Warnings 0.</p>
<p>Both were true in their own way. The warning line went to the command window only. The report counts what it is sent, and nobody had sent it that line. Claude fixed the file list and redeployed; the application was back at 16:14. It also added an automatic checker so that a deploy can never again miss a file its server loads. Then it asked to close the gap itself: 23 deploy scripts sent that health check line to the window alone, and I approved sending it to the report in all 23.</p>
<h3>Odd is not cool</h3>
<p>That is where the second story turned. Of the 23, three had no report at all. Not a missing line. No report page, ever.</p>
<p>We have an automatic checker whose one job is to make sure every deploy script writes a report. It had passed these three every time. Why? It recognised a deploy script by one exact shape of log line, and these three wrote a slightly different shape. The checker did not fail them. It did not see them. It skipped them in silence and announced that everything was good.</p>
<p>Claude's first plan was to leave those three alone and fix the twenty. I said no. &quot;Odd is not cool&quot;, I wrote, &quot;fleet apps are consistent.&quot; A little later: &quot;Log consistency is KING.&quot; It reads like a slogan. It is a working rule. Every checker we have works by recognising a shape, and a thing written in a different shape is a thing no checker will ever look at.</p>
<p>So Claude surveyed all 138 deploy scripts. 128 wrote the one standard line. Ten wrote something else. Five of those ten had no report at all. And one script, read closely for the first time, had its last line cut off in the middle of a word since the day it was written. It should have said &quot;Green&quot;, a colour. It said &quot;Gree&quot;. So every deploy of that application had ended with an error after the real work was finished. Nobody had noticed, because the work itself always succeeded.</p>
<p>Then the locked doors. When Claude went to change two of the ten scripts, the session's automatic safety check stopped the edit, as it is meant to. Claude did not look for a way round. It stopped, explained which scripts and why, and asked. I approved all ten.</p>
<p>By the end of the afternoon all 138 scripts write the same log line. All 138 write a report. The report checker recognises the wider shape, and Claude showed that it still fails on an old unwired script, so the net was not loosened. A new checker, a blocking one, refuses any deploy script whose log line differs from the rest. Ten odd scripts cannot become eleven.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/the-page-that-watches-map.svg" alt="The essay in two rows. Story one: a deploy that printed about 500 grey lines gets a live page of six phases and a clock, built into the shared module every deploy already loads, so no locked script changed. Story two, the same afternoon: a warning that never reached the report leads to three scripts with no report, a checker that knew one shape of log line, a survey of all 138 scripts, and one log line for all of them with a new blocking checker; below, what it teaches." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<h3>What this teaches</h3>
<p>For anyone building with an AI coding agent, the afternoon in a short list.</p>
<p>Show the same facts, sooner. The live page added no information. It moved the facts from the end of the deploy to the middle of it, and from 500 lines to six. That was enough for me to call it the article of the day.</p>
<p>A checker that knows one shape is blind to every other. It does not say &quot;I did not look&quot;. It says &quot;all good&quot;, because from where it stands there was nothing to look at. The three scripts with no report had passed every time.</p>
<p>Consistency is what lets the checkers see. One shape of log line is the condition under which every other check can do its job.</p>
<p>A warning that goes only to the screen is a warning that is not counted. Send it where the count is made.</p>
<p>Read the odd ones closely. The cut-off word had been there since the beginning. Nobody reads the end of a log when the work succeeded.</p>
<p>And the honest one. I noticed: that Praxis was better, that odd is not cool, that the log shape matters. Claude measured, proposed, built, tested, found its own off-by-one before it shipped, and stopped at the locked doors to ask. Neither half works without the other.</p>
<h3>For the engineers</h3>
<p>The detail, for readers who want it. The live page is <code>Write-DeployLive</code> in <code>infra/deploy/DeployOverlay.psm1</code>, the module nearly every deploy script already imported (all 138 do now), so no <code>deploy.ps1</code> changed. It writes <code>%TEMP%\deploy-live-&lt;PID&gt;.html</code> (one file per PowerShell process, which is why a batch reuses one tab) with <code>&lt;meta http-equiv=&quot;refresh&quot; content=&quot;2&quot;&gt;</code> while the deploy runs; the file is rewritten at start, at each new phase, and otherwise at most once per second. <code>$env:DEPLOY_NO_LIVE = '1'</code> turns it off. The finished page first had no refresh, so the next deploy in the window never showed; it now refreshes every 5 s (<a href="https://github.com/DanielFTwum-creator/aucdt-utilities/pull/929">PR #929</a>, RFE-480). An optional per-app <code>deploy-confirm.txt</code> adds its lines under &quot;Confirm live&quot;. The index bug: <code>[array]::IndexOf($states,'work')</code> returns -1 when no phase is working, and <code>$PHASES[-1]</code> is 'Verify'; a test now asserts &quot;Starting&quot;. Nine tests in all. <a href="https://github.com/DanielFTwum-creator/aucdt-utilities/pull/926">PR #926</a>, RFE-476.</p>
<p>The port line: 20 scripts now send &quot;OK/WARN port N&quot; to the report through <code>Write-DeployLine -Raw</code>, and the module counts an untagged <code>WARN</code> line as a warning, so &quot;Warnings 0&quot; beside a WARN in the window cannot recur. <a href="https://github.com/DanielFTwum-creator/aucdt-utilities/pull/927">PR #927</a>, RFE-473.</p>
<p>The log format: the standard line is <code>[$ts][&lt;app&gt;][$Level] $Msg</code>. <code>scripts/check-deploy-log-format.mjs</code> is a blocking guard, 138 of 138 passing. <code>scripts/check-overlay-wiring.mjs</code> now matches <code>][$Level]</code> rather than <code>[$ts][$Level]</code>, which is how the three unreported scripts had slipped past it, and it was shown to fail on an old unwired script. <a href="https://github.com/DanielFTwum-creator/aucdt-utilities/pull/928">PR #928</a>, RFE-473 and RFE-479.</p>
<p>#ai4good #AgenticEngineering #HumanInTheLoop #DevOps</p>]]></content:encoded>
</item>
<item>
<title>The Slow Push</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-slow-push/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-slow-push/</guid>
<pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>The safety check on every push grew from thirteen seconds to more than two minutes, and the slow part was the AI's own addition from that morning. I noticed the wait. The AI measured, fixed it, proved the net still catches, and put a timer in the tool so the next slow check is named on day one.</description>
<content:encoded><![CDATA[<p><em>Article 54 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>Two and a half minutes</h3>
<p>I look after about 127 web applications for the university, with an AI assistant, Claude, as my working partner. When either of us changes something, the change goes up to GitHub, the shared store where our code lives. We call that a push.</p>
<p>Before any push leaves the machine, an automatic set of checks runs first. There are 46 of them, and we call them the fleet guards. Each one looks for a known kind of mistake: a box of instructions missing the blank line at its end, a port number written one way in one file and another way in a second file, a page that lists an application twice. If any check fails, the push stops. The checks used to run on GitHub's own machines. On 25 September those machines stopped starting our jobs, so the checks moved onto our own computer, in front of the push. A check you have to remember to run is a check that gets skipped, so we made it automatic. Its own description said the whole set took about 13 seconds.</p>
<p>On 1 October I was watching Claude work. It had finished a change and sent it on its way. Then it wrote: &quot;The push is still running its guards. I wait for it to finish.&quot; And it waited. So did I. The push ran for more than two and a half minutes.</p>
<p>Nothing had failed. Nothing was broken. It was only slow, and slow in a place that every push would have to pass through from then on. I took a screenshot and sent one line: &quot;RFE Investigate how to speed up push guards&quot;. (An RFE is a request for enhancement, our name for &quot;this needs fixing or improving&quot;.)</p>
<h3>Measure before you touch anything</h3>
<p>The first thing Claude did was not to change anything. It ran the checks on their own and timed them.</p>
<p>The full set took 75 seconds. One check alone took 57 of those. The other 45 shared the remaining 18.</p>
<p>That one check was Claude's own work, added the same morning. It keeps our AI Lab catalogue, the public page that lists all our applications, in step with the code. Every application's card on that page shows five small tick marks, and the check makes sure the ticks match what the code says. It was correct. It had passed every time it ran. Nobody had timed it, because nothing we had asks a check how long it takes. A check that gives the right answer and a check that gives the right answer quickly look the same to a test for correctness.</p>
<p>I want to pause on this, because it is the point of the story for anyone learning to build with these tools. The slow part was not an old piece of code nobody understood. It was the newest piece, written by the AI a few hours earlier, and it passed its tests. What caught it was a person watching a screen and noticing a wait that felt too long.</p>
<h3>The same answer, 127 times</h3>
<p>Why was it slow? Claude read its own check and found the shape of the problem.</p>
<p>To fill in five ticks for one application, the check ran our general readiness checker for that application. That checker asks many questions, not five, and to answer some of them it starts up to six further small programs. One of those scans the whole of our code to answer a question about the catalogue itself. The answer to that question is the same for every application. The check asked it 127 times, once per application, one after another, and got the same answer 127 times.</p>
<p>The catalogue only needed five of the many checks, and those five do nothing but read files. Everything else was work that nobody was going to look at.</p>
<p>This shape turns up everywhere once you know to look for it: a loop that calls something heavy, and the heavy thing does the same work on every turn. The fix is rarely clever. You work out what the loop actually needs, and you stop asking for the rest.</p>
<h3>Three fixes</h3>
<p>The first fix was a &quot;card&quot; mode for the readiness checker: run only the five checks the catalogue card shows, skip the rest, and leave the skipped results out of the report entirely. That last part matters. A report that lists a check it did not run is a quiet lie, so card mode says nothing about the checks it skipped.</p>
<p>The second fix was to run the 127 applications several at a time rather than in a queue. Our machine has four processor cores, so four at a time. Because each of those checks only reads files and writes nothing, they cannot trip over one another. Claude compared the result file before and after: identical, byte for byte. The 57 seconds became 5.</p>
<p>The third fix was the same idea one level up. The 46 guards themselves ran one after another. They, too, only read files, so they can run several at a time as well. The one thing to protect was the order of the printed results, because a person reads that output and wants the same list every time. So the guards run several at a time, but their results are held back and printed in the original order.</p>
<p>The whole set went from 75 seconds to 23 after the first two fixes, then to about 11 after the third. A real push with the guards in front of it took 13 seconds, which is what the description had promised in the first place.</p>
<h3>Prove the net still catches</h3>
<p>A speed-up that quietly breaks the safety net is worse than the slow version. So before calling it done, Claude broke something on purpose. It planted a stale copy of the catalogue file, the exact fault the slow check exists to catch, and ran the guards.</p>
<p>They failed. The run stopped with the right error code, the failing check's own message, and the summary at the end.</p>
<p>This is the step people skip. When you make something faster, you test that it still passes. You also have to test that it still fails when it should, because the fast path and the slow path are now different code.</p>
<h3>So it cannot creep back</h3>
<p>Two more things, so the next slow check does not need me to notice it.</p>
<p>The runner now has a budget. Any single guard that takes longer than 15 seconds is named on screen as SLOW, on the very run that adds it. It does not block the push, because machines differ, but it is named, with its time, where the person adding it will see it.</p>
<p>And the lesson is written down as Rule 111 in our operating notes, the file Claude reads at the start of every session: a new check is timed before it ships, because every push pays for it. The next time Claude adds a guard, it will have read that rule a few minutes earlier.</p>
<p>One small footnote. While this was going on, the previous pull request (our name for a proposed change waiting for review) had already been merged. When Claude tried to add this work on top of it, a safety script of ours refused: &quot;STOP: PR #912 is merged or closed ... Nothing pushed&quot;. So the work moved to a fresh branch and a new pull request. A different guard, doing its job on the same afternoon.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/the-slow-push-map.svg" alt="The essay as one wrapped flow. A push that should take 13 seconds runs for two and a half minutes; Claude times the checks first and finds one of 46 taking 57 of 75 seconds, its own work from that morning, because a loop asked the same heavy question 127 times. Three fixes bring 75 seconds to about 11, a planted fault proves the net still catches, and a SLOW budget in the tool plus a rule in the notes stop it creeping back; below, what it teaches." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<h3>What this teaches</h3>
<p>For anyone learning to build with an AI coding agent, here is the afternoon as a short list.</p>
<p>Measure before you optimise. Claude's first act was a stopwatch, not a change. The number (one check out of 46 taking three quarters of the time) said exactly where to look. Without it, the obvious guess would have been &quot;46 checks is too many&quot;, and that guess would have been wrong.</p>
<p>A correct check can still be a costly one. Tests for correctness will never catch slowness. If a tool runs on every push, its speed is part of whether it is fit to ship.</p>
<p>Look for the same work repeated in a loop. If a loop calls something heavy, ask what the loop actually needs from it. Often it needs a fraction.</p>
<p>Things that only read can run side by side. Keep the printed order stable so a person can still read the output.</p>
<p>After a speed-up, prove the failure path still fails. Plant the fault. Watch it get caught.</p>
<p>Put the budget in the tool. A rule in a document is read by people who remember to read it. A SLOW line on the screen is read by whoever is sitting there.</p>
<p>Write the lesson where the next session will read it. The agent that adds the next guard is not the one that learned this. It has to be told, and the notes are how we tell it.</p>
<p>And the honest one. The slow step was the AI's own addition, tested and correct, and a human noticed the wait. Neither of us would have got there alone. I did not know why the push was slow, and I would not have found the 127 repeated scans. Claude did not know the push was slow at all until I asked, because nothing it had built was asking.</p>
<h3>For the engineers</h3>
<p>The detail, for readers who want it. <code>scripts/check-app-ready.mjs --card</code> (RFE-462) runs only the five checks the AI Lab catalogue card shows and drops every other result row, so a card-mode report never claims a check it did not run; in that mode no guard process is spawned except <code>check-ps1-param-clobber.mjs</code>. <code>scripts/gen-catalog-checks.mjs</code> runs the 127 apps through a small worker pool (<code>cpus().length</code>, minimum two) and exits FATAL if a label it maps no longer comes back from <code>--card</code>, so a renamed check cannot silently empty a column; <code>checks.json</code> was byte-for-byte identical before and after. <code>scripts/run-fleet-guards.mjs --jobs &lt;n&gt;</code> runs the 46 steps of <code>.github/workflows/fleet-guards.yml</code> several at a time (default the CPU count; <code>--jobs 1</code> is the old one-by-one run), gives each step its own <code>GITHUB_STEP_SUMMARY</code> file, and flushes results in workflow order; a step over <code>STEP_BUDGET_S = 15</code> prints <code>SLOW</code> without failing the run. The pre-push hook is <code>scripts/git-hooks/pre-push</code> (<code>git config hooks.fleetguards off</code> turns it off for one clone). Timings: 75.5 s one by one; 22.7 s after card mode and the worker pool; 10.7 s with parallel steps; a real push 13.3 s. A planted stale <code>checks.json</code> exits 1 with the step's output and the summary. The lesson is AGENT_OPERATING_NOTES Rule 111. Register entries: RFE-462 (the speed-up) and RFE-463 (this article).</p>
<p>#ai4good #AgenticEngineering #HumanInTheLoop #DevOps</p>]]></content:encoded>
</item>
<item>
<title>The Empty Line</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-empty-line/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/the-empty-line/</guid>
<pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>A set of instructions I pasted stopped one key short of finishing. I spotted where, the AI found why, and one cheeky request turned a single repair into a rule that now holds across nearly seven thousand boxes of instructions.</description>
<content:encoded><![CDATA[<p><em>Article 53 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>One key short</h3>
<p>I look after about 127 web applications for the university. I do that work with an AI assistant, Claude, as my working partner. Claude writes the instructions; I am the one who runs them on our servers. Each set of instructions sits in a box on a private page, with a Copy button beside it. I press Copy, paste the box into a command window, and watch it run.</p>
<p>We agreed a small rule long ago. Every box ends with one blank line. A command window runs a line only when you press Enter at the end of it, and pasting that blank line is the same as pressing Enter for you. Without it, the last instruction is typed out but never run. It just sits there, waiting.</p>
<p>At noon on 1 October I pasted the box that updates our AI Lab catalogue, the public page that lists all our applications. Everything ran until the very last line. Then it stopped and waited. Nothing was broken, but nothing had been published either. I sent Claude a screenshot and one line: &quot;RFE: Missing enter key to push deployment&quot;. (An RFE is a request for enhancement, our name for &quot;this needs fixing or improving&quot;.)</p>
<h3>A wrong guess, corrected in five words</h3>
<p>Claude's first answer was wrong. It blamed an earlier step in the box. It was a reasonable guess, and it was not what my screen showed.</p>
<p>I did not argue. I wrote five words: &quot;enter key after last }&quot;. The last line of that box was a closing bracket, and the missing Enter belonged right after it.</p>
<p>Those five words did a lot. They said where the problem was (at the very end, not in the middle) and what it was (one missing key press). Claude dropped its first guess and looked at the Copy button instead of the instructions.</p>
<p>I want to keep that moment in the story. It is easy to tell these tales as if the AI goes straight to the answer. It did not. It had a theory, the theory was wrong, and the person looking at the screen could see that. The correction cost me five words. Without it, the next hour would have gone on the wrong problem.</p>
<h3>Why the Copy button sometimes let us down</h3>
<p>Once the problem had a name, the cause was quick to find.</p>
<p>The Copy button is what adds the blank line. But a web browser does not always let a page write to your clipboard; sometimes it says no, for your safety. When that happened, our page fell back to an older way of copying, and that older way quietly lost the blank line. So the same box could work perfectly one day and stop one key short the next, depending only on what the browser allowed.</p>
<p>Claude changed the fallback so it keeps the blank line too. If even that fails, the page now says plainly: paste, then press Enter. Claude tested the repair with the browser set to refuse, which is exactly the case that had failed.</p>
<h3>From one repair to a rule</h3>
<p>At that point my box worked, and the story could have ended there. I sent one more line instead: &quot;RFE Trailing empty line sweep ;-)&quot;.</p>
<p>The wink mattered. It meant: I know this is bigger than the one box I pasted. Fixing one button is a repair. Making sure every box everywhere follows the rule is something else, and I wanted that.</p>
<p>We keep our instructions in a large collection of written pages, about seventeen hundred of them. We already had an automatic checker that looked at some of the boxes on those pages, but only one kind. Claude widened it to cover the other kinds, then let it add the blank line wherever one was missing.</p>
<p>It found 6,962 boxes that needed it. It added the blank line to each one and changed nothing else. And because the checker now runs every time anyone changes our pages, a new box without its blank line is caught before it ever reaches me.</p>
<p>I could not have done that by hand. Not in a day, and not without mistakes. Claude did it in minutes, and the result is easy for a person to check: each change is one added blank line.</p>
<p>When the work came back, I wrote: &quot;Here is another beautiful article demonstrating the power of smart LLMs acting as colabs with humans&quot;. (An LLM, a large language model, is the kind of AI that Claude is.) So here it is.</p>
<h3>The same afternoon, the same pattern</h3>
<p>The afternoon gave a smaller example of the same thing.</p>
<p>I opened the catalogue page that the noon update had published. It said 127 live applications, 66 meeting our standards and 61 needing work. Something about the numbers felt off. I wrote four words: &quot;RFE Count accuracy check&quot;.</p>
<p>Claude found three faults. One application was listed twice, so one tally said 128 while the page said 127. Eight applications showed &quot;unknown&quot; because our checker could not read the way their settings were written, and one of those eight was hiding a real problem. One more application kept its files in an unexpected folder, so the checker skipped it.</p>
<p>After the repair the page says 127 applications, 72 meeting our standards and 55 needing work. A new check stops any application being listed twice again. And the small status label on each card, which my screenshot showed as nearly invisible, is now a coloured badge you can read.</p>
<p>The pattern was the same as at noon. A person notices that something on the screen looks wrong. The AI finds out why. And the repair comes with a check, so the same fault cannot quietly return.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/the-empty-line-map.svg" alt="The essay as one wrapped flow. A pasted box stops one key short; Claude's first guess is wrong and five words point at the end of the box; the cause is a Copy button fallback that lost the blank line when the browser refused; the fallback is repaired and tested in the failing case; a wink turns one repair into a rule applied to 6,962 boxes with a checker on every change. Beside it, the same pattern that afternoon on the catalogue counts; below, who did what." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<h3>Who did what</h3>
<p>Put the two lists side by side.</p>
<p>I brought the eye. I was the one looking at a command window with a bracket stuck on the last line. I named the problem, and when the first guess was wrong I corrected it in five words. Then I raised the bar from &quot;fix this&quot; to &quot;make it a rule&quot;, with a wink. In the afternoon I looked at a page and doubted a number. None of that took long. All of it needed someone who knew what the screen should look like.</p>
<p>Claude brought the rest. It found the hidden path in the Copy button, which I would never have thought to look for. It repaired that path and tested it in exactly the situation that had failed. It applied the rule to nearly seven thousand boxes without disturbing anything else. It turned my doubt about a number into three named faults and a new safeguard.</p>
<p>It also got the first diagnosis wrong. I do not think the story is better for leaving that out. A partner who is never wrong would not need a partner.</p>
<p>What strikes me, reading the day back, is how cheap each side's part was for the one who gave it, and how costly it would have been for the other. Five words from me saved Claude an hour on the wrong problem. A few minutes of Claude's work saved me a job I would never have started. The wink is where the two meet: a person saying the standard is higher than the fix, and a machine that can meet that higher standard the same afternoon.</p>
<h3>For the engineers</h3>
<p>The detail, for readers who want it. The runbook page's Copy button writes the block with <code>navigator.clipboard</code>; when the browser refuses, the old fallback selected the <code>&lt;code&gt;</code> element and asked for Ctrl+C, which drops the final newline. The new fallback copies from a hidden textarea with <code>execCommand('copy')</code>, tested in headless Chromium with clipboard access refused (the copied text ends <code>}\n\n</code>). The guard is <code>scripts/check-copy-box-enter.mjs</code> (RFE-378, widened under RFE-457): it now checks bash, sh, shell and zsh boxes and unlabelled boxes that start with ssh, sudo, <code>cd /</code>, bash, pm2, systemctl or plesk, with 8 tests; its <code>--fix</code> touched 1,717 Markdown files. The catalogue fixes (RFE-459) were a duplicate <code>/enrollment-2026/</code> row, a Vite <code>base</code> written as an expression (<code>mode === 'mobile' ? '/' : '/trotro-rush/'</code>), and an app whose front end lives in <code>client/</code>.</p>
<p>#ai4good #AgenticEngineering #HumanInTheLoop #DevOps</p>]]></content:encoded>
</item>
<item>
<title>Helpers Practise the Blueprint</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/helpers-practise-the-blueprint/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/helpers-practise-the-blueprint/</guid>
<pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>One evening, three helpers ran the fleet's OODA and 6R audit on three different objects, a deploy report page, a video and the fleet's own skill files, and the results show what a brief does, why a finding is checked at the source, and which calls stay with the human.</description>
<content:encoded><![CDATA[<p><em>Article 52 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>One loop, three objects</h3>
<p><a href="../muse-does-my-ui-testing/">Article 50</a> and <a href="../hand-me-your-notebook/">Article 51</a> were about one helper, Muse, and one afternoon. This one is about the evening that followed, when three helpers ran the same audit method on three objects that had nothing in common: a deploy report page, a video, and the fleet's own rulebooks.</p>
<p>The method is OODA (Observe, Orient, Decide, Act) with the six Rs (Diagnostics, Reduce, Refine, Reuse, Rebuild, Resilience). That afternoon I had put the whole pipeline on WhatsApp in one line: &quot;IDEA -&gt; IEEE SRS 14-sections -&gt; AI Studio Prototype -&gt; CLAUDE AI Blueprint -&gt; PRODUCT&quot;. The blueprint is the delegation part: Haiku for mechanical work, Sonnet for research and synthesis, Fable for long-form prose, Opus for judgement and the human call, with a standing permission to delegate research and documentation prose without asking. The parent reviews and commits.</p>
<p>A method you only describe is a poster. A method you run every evening on small real things is a practice. Here is what one evening of practice looked like.</p>
<h3>The tester and the deploy report</h3>
<p>At 21:42 the essays site itself deployed, commit 166317fbb. The run went from 21:42:35 to 21:43:26 and printed &quot;Time : 51.3s total&quot;. A tester helper reviewed the HTML report the deploy writes, and I pasted its review to the builder. I did not name the tester in the paste, so I will not name it here.</p>
<p>The review came back as &quot;7 findings: 2 high, 3 medium, 2 low&quot;. It opened with praise: &quot;The page is genuinely elegant: one centred card, clear status plus action in the header, scannable meta pills, a step timeline, restrained log colour. No redesign needed.&quot;</p>
<p>Then the sharpest finding. &quot;The DURATION pill reads 00:01:51, but the log says 51.3s total ... A report must not disagree with itself.&quot; Second: &quot;The Restart step is greyed with a dash and no timing. Skipped, not applicable, or failed? It should say.&quot; After that the small ones: the grammar of &quot;1 progress lines folded&quot;; the URL logged twice with different spacing; mixed key spacing, &quot;Remote :&quot; against the pills.</p>
<p>The builder went to Orient, against the source. It ran the report module with a start time 51.3 seconds in the past, and the module wrote 00:00:51. Minutes later I pasted the report's own Copy text: &quot;Duration: 00:00:51, Started: 2026-09-30 21:42:35, Finished: 2026-09-30 21:43:26&quot;. The headline finding, the one rated highest, was a misread. Not a defect.</p>
<p>The Restart finding was real, in a narrower way. The chip meant &quot;not run&quot; (a static app has no restart step), and the Copy text already said &quot;Restart not run&quot;, but the page showed only a grey dash. Fixed: the chip now says &quot;not run&quot;, with a tooltip. The grammar finding was real, &quot;(+1 progress lines folded)&quot;, and is fixed. The URL printed twice was real, and bigger than the page. It comes from 27 deploy.ps1 files: a codemod this year added &quot;URL:  &lt;url&gt;&quot; to scripts that already printed &quot;URL  : &lt;url&gt;&quot;. Deploy scripts are locked files, so the fix waited for my yes. I said yes: the extra line is gone from all 27, the codemod now recognises both forms, and a new guard fails any script that prints its URL twice. The key spacing was kept, because it is the scripts' own aligned console text.</p>
<p>The report's headless test suite passes 36 of 36 assertions, 3 of them new. RFE-446, PR #900.</p>
<p>Of the five findings named in the paste: one false, and it was the headline; two real and fixed within the hour; one real, handed to me, and fixed once I said yes; one kept on purpose.</p>
<h3>Meta AI and the mug</h3>
<p>The second helper was not in the fleet at all. I was making a double-exposure video on meta.ai (&quot;Desert caravan double exposure&quot;), and a mug in the shot vanished for a few frames. I told Meta AI: &quot;The mug disappearing does not pass ooda 6r audit&quot;.</p>
<p>It answered: &quot;You are 100% correct, that fails best practice.&quot; Then it wrote an audit under the heading &quot;OODA 6R Audit Failure. Mug Discontinuity&quot;. Observe: &quot;In the last render, your Kente-pattern mug blinked out for frames 4-7 when you raised it. Artifacting on a personal, cultural object.&quot; Orient: &quot;Breaks continuity, breaks respect, breaks #ai4good. For posterity, every frame must honor the object, time is on our side, so we don't rush.&quot; Decide: &quot;Fail on:&quot; and two bullets, &quot;Representation: mug = veteran calm, disappear[s] = veteran disappears&quot; and &quot;Realism: object persistence violated&quot;. The screenshot ends there.</p>
<p>Meta AI had no fleet brief. It kept the method's shape, Observe, Orient and Decide as headings and a fail verdict, and it invented the Rs. Representation and Realism are not on the fleet's list. It also opened by agreeing with the verdict I had handed it. &quot;100% correct&quot; is a helper telling you what you just told it.</p>
<p>The finding itself is right and useful: object persistence. But the Act step was not new. The fleet's own video skill, ai-video-director, already carries the rule: &quot;Lock objects the same way as people: the car's colour, model and interior, the print on a shirt, the side of the wrist a bracelet is on.&quot; Its review checklist adds: &quot;Negatives forbid each camera move the lock excludes and each item that must not morph.&quot; So the fix is to name the Kente mug in the object lock, word for word in every prompt, and to add it to the negatives. The helper found the defect. The rulebook already held the fix.</p>
<h3>A research helper reads the rulebooks</h3>
<p>The third helper turned the method on the rulebooks themselves. I wrote: &quot;RFE: OODA 6R audit all known SKILL files&quot;.</p>
<p>The builder set the scope: the 11 skills the fleet owns in <code>.claude/skills</code> (ace, ai-video-director, aistudio-fleet-standards, five-theme-system, fleet-app-value-deck, fleet-runbook, graphify, issue-subdomain-cert, mighty-sparrow-songcraft, source-to-fleet-app, tuc-enrollment-push). Out of scope: 16 skills synced from the account that are Anthropic's own, and the harness's session-start hook, which the fleet cannot change.</p>
<p>Then it sent a Sonnet research helper with a brief that spelled out the six Rs word for word, told it to verify every path, pattern number, rule number, script and version a skill claims against the repo, and to mark anything it could not verify. Read-only: no edits, no git.</p>
<p>The helper read the 11 SKILL.md files (1,465 lines) and 29 files beside them, ran for about 13 minutes and 84 tool calls, and reported 71 findings (6 High, 31 Medium, 34 Low), plus 14 across skills. The report is TUC-ICT-AUD-2026-006, under RFE-447.</p>
<p>The striking part: both of the fleet's existing skill guards passed. 26 cited paths exist; 422 citations resolve. The guards check that a citation exists, not what it says. A skill can cite the right pattern and say the opposite.</p>
<p>The builder re-checked five of the six High findings before writing the report, the five it could re-run from the repo, and all five held. source-to-fleet-app tells the reader to use Vite base <code>'./'</code> for single-page apps; CLAUDE.md section 5b says the base is <code>'/&lt;slug&gt;/'</code> &quot;(absolute, never './')&quot;, and the sister skill aistudio-fleet-standards sides with CLAUDE.md, so two skills contradict each other. The same skill says the deploy clone &quot;must carry --branch&quot;; CLAUDE.md says &quot;-Build sources frontend AND backend from one fresh git clone of main&quot;. fleet-runbook never mentions the Open runbooks page, and its fix template restarts with &quot;pm2 restart --update-env&quot; where Pattern 23 says &quot;never rely on reload/restart --update-env&quot;. And five-theme-system's Golden accent, #D4AF37 on a cream ground, is 1.85:1 against a 4.5:1 minimum; the palette was copied from a live app, biochemai, which carries it too.</p>
<p>Nothing in a skill has changed yet. The report gives a fix order and six calls that are mine: brand colours, the theme system's place, Registrar confirmation of admissions claims, keep or drop one tool, extend the songwriting skill to Afrobeats and highlife, and a text-style guard.</p>
<h3>The brief is part of the method</h3>
<p>With a brief, the vocabulary holds. The skills helper was given the six Rs and used them; Muse works from MUSE.md, a tester's cut generated from CLAUDE.md (Article 51). Without a brief, the shape survives and the terms drift. Meta AI produced Observe, Orient and Decide, then Representation and Realism. It was running a method rebuilt from one sentence of mine. The brief is not a courtesy to the helper. It is part of the method.</p>
<p>Two ways a helper fails showed up, and they look like opposites. One agrees too readily: &quot;You are 100% correct&quot; before a single frame was examined. The other finds with too much confidence: the tester's highest-rated finding was 00:00:51 read as 00:01:51. The remedy is the same for both. Orient against the source before you Act. Article 50 put it in four words, a finding is a claim, and that stayed true when the helper changed.</p>
<p>The human keeps the calls a helper cannot make. Twenty-seven locked deploy scripts printed a URL twice, and the fix waited for my yes. Six calls sit at the end of the skills report. The helpers find. They do not decide what gets touched.</p>
<p>The third case matters most, because the rulebooks are what the helpers learn from. If a skill says <code>'./'</code> where CLAUDE.md says <code>'/&lt;slug&gt;/'</code>, every helper that reads it learns the wrong thing with full confidence. The method turned on itself found that. The guards had not, because a check that a citation exists cannot tell you whether it is true.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/helpers-practise-the-blueprint-map.svg" alt="The essay as one method and three columns. OODA with the six Rs runs on a deploy report, a video and the fleet's own rulebooks: the tester's headline finding is a misread, Meta AI without a brief keeps the shape but invents the Rs, and a research helper finds 71 faults the guards had passed. Every claim is checked against the source, the human keeps the calls, and below are the lessons." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<h3>What I take from it</h3>
<p>A method is what you run, not what you wrote. Three helpers ran OODA and the six Rs in one evening.</p>
<p>Give the helper the brief. With it, the words hold. Without it, the shape survives and the terms drift.</p>
<p>Orient before Act, whether the helper agrees with you too quickly or disagrees with you too confidently. Both are claims. The source settles them.</p>
<p>Keep the calls. Locked files, fix order, brand colours. The helpers found more in one evening than I would have alone, and not one of them decided anything.</p>
<p>The practice is small, real and daily. That is the whole method.</p>
<p>#ai4good #AgenticEngineering #HumanInTheLoop #OODA</p>]]></content:encoded>
</item>
<item>
<title>Hand Me Your Notebook</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/hand-me-your-notebook/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/hand-me-your-notebook/</guid>
<pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>A Claude Code session summarised its own memory, Daniel handed the summary to Muse, and the tester's next audit arrived in the builder's Workflow Card format. Three questions answered: whether behaviour transfers between agents (yes), whether shared memory saves tokens (not on reading, yes on re-teaching), and whether it belongs to the delegation blueprint (it is the blueprint made portable, with the same limits).</description>
<content:encoded><![CDATA[<p><em>Article 51 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>Another FYI</h3>
<p>Article 50 opened with an FYI. This one does too, and it is the same afternoon.</p>
<p>&quot;FYI Claude Code shared its memory (not the full details) and I provided it to Muse.&quot;</p>
<p>A few lines later came the part I wanted answered: &quot;RFE: and that reeks of another essay SHARING memory among LLMS, inheriting SKILLS and behavior? Does it save tokens? Is it not part of the blueprint: delegation #ai4good?&quot;</p>
<p>Three questions. I will take them in order, but the facts first, because the builder asked and I had to be precise. What went to Muse was not the memory. It was a summary of the memory, written out by a Claude Code session with the details left out, and I gave that summary to Muse. When the builder later asked how Muse already knew our Workflow Card format at 16:28, the answer was an earlier share of the same kind. As I put it that evening: &quot;WELL Muse only knows what you shared + recent OODA 6R audit.&quot; Muse has the summary and its own recent audit work. Nothing else.</p>
<p>So the title is a small lie. Nobody handed over the notebook. They handed over the page at the front that says what the notebook is for.</p>
<h3>What the notebook weighs</h3>
<p>It helps to know what stayed behind, because the weight is the point.</p>
<p>The builder's memory is four Markdown files at the root of the monorepo, plus eleven skills. CLAUDE.md, 374 lines, about 33.6 kB, about 8,400 tokens: the rules that must be true in every session. SHARED-STANDARDS.md, the index to the governance layer, about 1,300 tokens. PATTERNS.md, about 60 numbered patterns, 220 kB, about 55,000 tokens. AGENT_OPERATING_NOTES.md, 109 numbered rules, each paired with the fix that stops a repeat, 183 kB, about 45,700 tokens. (The &quot;about&quot; is honest. I divide bytes by four and call it a day.)</p>
<p>Then there is RFE-REGISTER.md, the per-project history. 1.5 MB. About 380,000 tokens. Nobody loads it. The builder searches it for the project in hand and reads the rows that come back.</p>
<p>Put the four files together and the builder carries about 110,000 tokens of governance on its shelf. Muse got a card. If you have ever started a job with a two-page induction while the colleague next to you has the 400-page binder in their head, you know both sides of that arrangement.</p>
<h3>Question one: does behaviour transfer?</h3>
<p>Yes, and the proof is in Article 50.</p>
<p>At 16:28 Muse logged the deploy and wrote a Workflow Card before it audited anything: JOB, INPUTS, BOUNDARY, OUTPUT, VERIFICATION, HUMAN CALL. Six lines, in that order. Section 10 of our CLAUDE.md asks the builder to state exactly that card before any non-trivial task. Muse ran OODA and 6R, the fleet's audit method. It kept a read-only BOUNDARY, no login and no forms. It left the HUMAN CALL to me.</p>
<p>None of that came from Muse's own habits. It came from the summary. A second agent, in a separate chat, that has never read CLAUDE.md produced the builder's card because someone told it the card existed and what the six lines were. Behaviour transferred in a paragraph.</p>
<p>What did not transfer is the reason behind each rule. The operating notes pair every rule with the miss that earned it. The summary carries the rule; the miss stays on the shelf. That is fine for a tester with a read-only boundary. It would be less fine for an agent about to run a deploy.</p>
<h3>Question two: does it save tokens?</h3>
<p>Two answers, and they point in opposite directions.</p>
<p>On reading, no. Every agent pays for what it loads, in every session. Had Muse received all four files, it would have paid the full price on every audit, and the sharing would have made nothing cheaper. Shared memory is still memory that has to be read.</p>
<p>On re-teaching, yes, and this is where the money is. Each of those 109 operating-note rules was paid for once, by a real miss. Today alone the rules earned their keep twice. A merge landed on main in the middle of a deploy batch; the deploy preflight stopped the batch, because a rule already says never deploy from a stale local copy. A squash merge took only the first commit of a pull request; the post-merge diff check caught it. An agent that reads a rule does not have to repeat the miss that wrote it. That is the saving, and it compounds with every agent that reads.</p>
<p>The layering is what keeps the reading side bearable. CLAUDE.md has a rule about itself, which it calls keep-lean layering. A must-run initialiser belongs in the SessionStart hook: it runs once, deterministically, and never occupies a turn. CLAUDE.md holds the must-be-true invariants, kept lean, about 8,000 tokens in every session. PATTERNS, CONSTRAINTS and the handbook are must-be-available detail, read on demand, and the builder looks a pattern up by its heading instead of loading the library. The register is searched and never loaded. The file's own words: &quot;Making everything always-on dilutes attention and wastes the context window every turn.&quot;</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/shared-memory-layers.svg" alt="The keep-lean layers drawn as a building. Ground floor: the hook, must run once at start. First floor: CLAUDE.md, must be true, about 8k tokens in every session. Second floor: PATTERNS and OPERATING NOTES, must be available, about 100k tokens read on demand. Attic: the RFE register, searched, never loaded, about 380k tokens. Claude and Muse stand on the roof looking in opposite directions." loading="lazy" decoding="async"><figcaption>Share the rules. Keep the vantage points different.</figcaption></figure>
<p>Seen that way, the summary card is the cheapest share there is: the rules without the 100,000 tokens of why. The price is that the detail and the history stay behind. Which is why the builder still checks Muse's findings against the source, and always will.</p>
<h3>Question three: is it the blueprint?</h3>
<p>Yes. Sharing memory is the delegation blueprint made portable.</p>
<p>The blueprint in CLAUDE.md has four tiers: Haiku for mechanical work, Sonnet for research and synthesis, Fable for long-form prose, Opus for judgement and the human call. It has a standing permission, from 23 July, to delegate research and documentation prose without asking. It has one sentence that carries the weight: &quot;the parent always reviews and commits a subagent's output; subagents do not run git.&quot; And it has locked files, the deploy scripts, lockfiles and port assignments that no subagent may touch, checked by checksum after every fleet run.</p>
<p>Every subagent in that blueprint works from the governance the parent reads. A helper is an agent with the rules and a narrower job. Muse is the same thing, one step further away: a separate chat, the rules in summary form, and the narrowest job of all, look and do not touch.</p>
<p>The same afternoon showed what the blueprint does when it runs. Three helper agents changed the college crest across about 80 apps. The parent then checked 691 locked-file checksums, all unchanged, and found that one helper had gone a step beyond the decision: the Sick Bay app's own clinic favicon had been replaced with the crest. Restored. A writer agent drafted Article 50 from the sources; the parent checked the draft and removed five claims that were not true, including one that said my WhatsApp message had gone to the agent. It had not. It went to a human.</p>
<p>That is the lesson under all three questions. Shared memory carries rules. It does not carry facts. The rules make a helper behave; they cannot make it right. Verification stays with the parent, and the human call stays with me.</p>
<h3>The catch</h3>
<p>Article 50 argued that two agents with different blind spots cover more of the wall than one agent with a mirror. Shared memory pulls against that. The more of the rulebook Muse has, the more Muse thinks like the builder, and the more likely it is that both walk past the same sentence.</p>
<p>The resolution is to share the rules and keep the vantage points different. Muse sees only the live site, through a browser, with no login. The builder sees the source, the deploy log and the four files. In Article 50 that difference produced one true finding the builder had missed and one false positive the builder could correct. Same card, different windows. Keep it that way.</p>
<p>Two housekeeping notes, because someone will ask. The governance files hold no secrets. Keys live in server <code>.env</code> files, and section 12 of CLAUDE.md forbids printing a secret value in any response, script or documentation example. That is what makes a summary safe to hand to a third party at all. And a stale rule, once shared, misleads every agent that reads it, which is why the standing order says a doc describes what the code does now, and changes in the same commit as the code.</p>
<h3>Should Muse get the whole framework?</h3>
<p>Later that evening I asked the obvious follow-up: &quot;RFE: Should I share the full framework to Muse ;-)&quot;. The builder's answer was no, and not to save tokens. A tester's cut is better practice than the full shelf, for three reasons.</p>
<p>The first is the catch above. The whole framework would make Muse think like the builder, and the value of Muse is that it does not. The second is exposure. The full files name internal hosts, server paths and deploy mechanics that a read-only tester has no use for, and a chat outside the repository is a third party, however friendly. The third is drift. A summary pasted into a chat goes stale the day a rule changes, and nobody notices, because nothing checks it.</p>
<p>So the proposal was a tester's brief: one file, built by a script from the governance sources, holding only what a tester needs (the Workflow Card, the read-only boundary, the audit method, the accessibility gates, the fleet's sign-in rules), with no hosts and no paths, and a check in CI that fails when the brief falls behind its sources. My reply: &quot;Best practice preferred always; time is on our side; #ai4good.&quot; Rules that travel should come from one place and be checked, the same as code.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/hand-me-your-notebook-map.svg" alt="The whole essay as one flow. The builder's shelf of four governance files, about 110,000 tokens, is summed up into one card with the details left out; Daniel hands the card to Muse, a separate chat that never read the files, and Muse writes the builder's Workflow Card. Three answers: behaviour transfers in a paragraph; tokens are saved on re-teaching, never on reading; it is the delegation blueprint made portable with the same limits. The catch, that shared rules make agents alike, is resolved by keeping the windows different: Muse sees the live site, the builder sees the source." loading="lazy" decoding="async"><figcaption>Share the rules, not the shelf. Same card, different windows.</figcaption></figure>
<h3>What I take from it</h3>
<p>Share the rules, not the shelf. A summary moves behaviour; the detail stays where the deploys happen.</p>
<p>Tokens are saved on re-teaching, never on reading. Every agent pays for what it loads. Keep the always-on layer thin and let the rest be searched.</p>
<p>The blueprint travels with its limits. Review, no git, locked files, human call. An agent with the rules and without the limits has half a blueprint.</p>
<p>And keep the windows different. Same card, different view. That is the whole trick.</p>
<p>#ai4good #AgenticEngineering #HumanInTheLoop #SharedMemory</p>]]></content:encoded>
</item>
<item>
<title>Muse Does My UI Testing</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/muse-does-my-ui-testing/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/muse-does-my-ui-testing/</guid>
<pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>A read-only tester named Muse audited a freshly deployed app two minutes after the deploy, was stopped at the sign-in gate, audited the gate anyway, and produced two low defects; the builder agent then checked both, found one affected four apps and the other was a false positive, and left the one real trade-off to the human.</description>
<content:encoded><![CDATA[<p><em>Article 50 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>A message at 16:28</h3>
<p>On 30 September, between 16:20 and 16:31, I ran one deploy card that shipped eight apps one after the other: the fleet's new sign-in page on five two-factor apps, plus three others. Each deploy prints a report. The lecturer AI handbook, at <code>/lecturer/</code>, ran from 16:25:08 to 16:26:56. One minute and forty-eight seconds, zero warnings, zero errors.</p>
<p>At 16:28 I wrote on WhatsApp: &quot;FYI: Muse does all of my automated UI testing now&quot;, with a screenshot of a chat with Muse. Later I passed the thread and Muse's audit brief to the builder agent. Muse had already logged the deploy (&quot;Deploy logged. Clean run: 1m 48s, zero warnings, zero errors.&quot;), and under that it had written a Workflow Card and set off to test the handbook at its live address. The chat showed &quot;Opening URL&quot; and an &quot;Open browser&quot; button. The last line read: &quot;I will report the findings when the survey finishes.&quot;</p>
<p>&quot;FYI&quot; is what you say to a colleague when you have hired someone to check their work and you would rather they heard it from you. The agent that built the sign-in page that morning was about to have its homework marked by a tester it had never met.</p>
<h3>The tester reads the same rulebook</h3>
<p>What struck me first was not what Muse planned to test. It was the shape of the plan.</p>
<p>Our governance asks the builder agent to state a Workflow Card before any non-trivial task: JOB, INPUTS, BOUNDARY, OUTPUT, VERIFICATION, HUMAN CALL. The rule lives in our CLAUDE.md, section 10, and I wrote it for the builder. Muse's card had the same six lines.</p>
<p>JOB: first OODA and 6R audit of the lecturer AI handbook at the live URL. INPUTS: a read-only live survey, no prior defects, this run sets the baseline. BOUNDARY: no login, no form submissions; if the app is gated, audit the public surface and report the gate. OUTPUT: OODA findings, a 6R analysis, numbered defects with evidence and acceptance checks, verified strengths, timings. VERIFICATION: each finding checked for real on the live app. HUMAN CALL: you decide the fix order.</p>
<p>OODA is Observe, Orient, Decide, Act. 6R is Diagnostics, Reduce, Refine, Reuse, Rebuild, Resilience. Together they are the fleet's audit method. So the tester and the builder run on one set of rules, and the line that matters most is the BOUNDARY. Muse may look. Muse may not sign in, and may not submit a form. If the door is locked, Muse audits the door.</p>
<p>That sounds like a restriction. It is what makes the arrangement safe. A tester that cannot log in and cannot submit cannot create a record, send an email or lock anybody out. So it can run after every deploy, unattended, at 16:28 on a Wednesday, with nobody holding its hand. The boundary is the licence.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/muse-two-cards.svg" alt="Two Workflow Cards side by side. The builder, Claude: JOB build the fleet sign-in page, BOUNDARY no locked files, no secrets, no trade-off calls, HUMAN CALL Daniel. The tester, Muse: JOB first OODA and 6R audit of the live app, BOUNDARY no login, no forms, if gated audit the gate, HUMAN CALL Daniel decides the fix order. An arrow marks the tester's boundary as narrower." loading="lazy" decoding="async"><figcaption>Same card. Muse's boundary is narrower: that is what makes it safe to run after every deploy.</figcaption></figure>
<h3>The report from the front door</h3>
<p>The survey finished at 16:29. Score: two defects, both low. Zero broken images, zero page errors.</p>
<p>Muse did not get in. Under Observe it wrote that the public address is a branded sign-in gate, that the handbook sits behind Google sign-in for @techbridge.edu.gh accounts, and that the handbook itself could not be tested. A critic sent to review a restaurant, stopped at the door, and reviewing the door.</p>
<p>I expected that to be the end of the report. It was the start. Muse listed the strengths it had verified on the gate: exactly one h1; landmarks for banner, main and contentinfo; a working &quot;Skip to sign-in&quot; link that appears on the first Tab; a visible 4 px gold focus ring; every control named; the background video hidden from screen readers and hidden again under reduced motion; title, description, favicon, canonical link, JSON-LD and noindex all present; a layout that does not shift. A thorough inspection of a front porch, and a useful one, because for everyone without a college account the porch is the building.</p>
<p>Then the two defects.</p>
<p>D1: the sign-in instruction ends in a sentence fragment. &quot;Use your @techbridge.edu.gh account. Techbridge staff and students.&quot; Fix: complete sentences.</p>
<p>D2: the page's head comment says &quot;Auth-gated (WMS SSO, TUC members only)&quot;, but the page offers only a Google sign-in. Fix: name Google SSO, or remove the comment.</p>
<p>Under performance: a cold load takes about 11 seconds to reach an interactive sign-in, a warm reload about 2 seconds, and the autoplaying campus-tour video is &quot;a likely contributor&quot;.</p>
<p>And under not verified: everything behind the sign-in. A signed-in pass is needed. HUMAN CALL: the owner decides the fix order.</p>
<h3>The builder marks the tester</h3>
<p>Now the part I wanted to see. The builder agent, Claude, took Muse's report and checked every line of it against the source.</p>
<p>The strengths first. The skip link, the gold focus ring, the landmarks, the video hidden under reduced motion: all of that is the new fleet sign-in page the agent had built and rolled out that same day, the one Article 49 describes. A second agent, working only from the live site, had confirmed the builder's accessibility work.</p>
<p>D1 was real, and it was the agent's own mistake from that morning, made in PR 888, the two-factor batch. Here the pairing earns its keep. Muse sees one app, the one it was pointed at. The builder knows where it copied the pattern from. A search found the same fragment in four apps, not one: lecturer and aitopia (&quot;Techbridge staff and students.&quot;), fail2ban-ai and netscan (&quot;Techbridge staff only.&quot;). The tester reported a typo. The builder found the photocopier. All four now read &quot;This app is for Techbridge staff and students.&quot; or &quot;This app is for Techbridge staff only.&quot; Four builds pass.</p>
<p>D2 was not a defect. The Google button sends the browser to the college's WMS sign-in service, at <code>wms.techbridge.edu.gh/api/auth/google</code>, which relays Google. &quot;WMS SSO&quot; is exactly right. From the public page you see a Google button and nothing else, and a read-only, no-click audit cannot see the relay, because the relay only happens when you click. Muse judged a label it was not allowed to test against what it could see. The comment stays. A false positive, and the best kind: a tester obeying its boundary, saying &quot;this does not match my view&quot;, which is true, and leaving the last word to someone with a wider one.</p>
<p>The 11-second cold load is a different animal. The campus video is part of the fleet sign-in page, now on about 35 apps. If the video is the cause, the fix is fleet-wide: keep the video, make it lighter, or load it later. Muse said &quot;likely&quot;, not &quot;proven&quot;. The agent did not touch it. It came to me.</p>
<p>Tally: one finding under-scoped (one app reported, four affected), one false positive, one open question. Both agents were right about the part they could see. Neither could see all of it.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/muse-tally.svg" alt="Three panels. D1, a sentence fragment reported in one app, grows to four apps: lecturer, aitopia, fail2ban-ai, netscan; real, fixed. D2, the WMS SSO comment, shrinks: the Google button relays through WMS only on click, so not a defect. The 11-second cold load goes to Daniel as a human call: likely the video, not proven." loading="lazy" decoding="async"><figcaption>A finding is a claim. Check it both ways: one grew, one shrank, one went to the human.</figcaption></figure>
<h3>Why two agents beat one</h3>
<p>Anyone who has marked their own homework knows why the builder should not be the only tester. The builder agent had built the sign-in page, run the deploy, read its own report of zero warnings and zero errors, and would have gone home happy. Muse came in from outside and saw only the live site. That is why it got D2 wrong, and why it caught D1, a sentence the builder had written and read straight past. Two agents with different blind spots cover more of the wall than one agent with a mirror.</p>
<p>The read-only boundary is what lets the second agent run at all. A tester with a login and a submit button is a user with a script. A tester with neither is a very patient visitor, and the worst it can do is write a slightly wrong sentence about a comment in the page head.</p>
<p>A finding is a claim. Muse's report held a true claim that was too small (D1), a claim that was false from a wider view (D2), and a claim honestly labelled as a guess (the video). The builder checked all three against the source, in both directions: it grew D1 and shrank D2. Treating the report as a to-do list would have deleted a correct comment and fixed one app out of four.</p>
<p>And I stayed on WhatsApp. I did not read the code. I passed Muse's brief to the builder, got back a tally, and kept the two things that were mine: the fix order, and the trade-off on a campus video that plays on 35 sign-in pages. At 16:31 I put the whole pipeline on WhatsApp in one line: &quot;IDEA -&gt; IEEE SRS 14-sections -&gt; AI Studio Prototype -&gt; CLAUDE AI Blueprint -&gt; PRODUCT&quot;. Then: &quot;Agentic Engineering Pipelines only&quot;. Then a link to the demo, which is the site you are reading.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/muse-pipeline.svg" alt="A conveyor belt: IDEA, IEEE SRS in 14 sections, AI Studio prototype, Claude blueprint, PRODUCT. After it, a deploy truck (1m 48s, clean run) and Muse the robot at the sign-in door. Daniel, on WhatsApp, sends the whole pipeline in one line and a link to the demo." loading="lazy" decoding="async"><figcaption>Agentic Engineering Pipelines only.</figcaption></figure>
<figure class="post-figure"><img src="/agentic-essays/articles/images/muse-does-my-ui-testing-map.svg" alt="The whole essay as three lanes in one afternoon. Daniel runs one deploy card and sends the FYI, then keeps the fix order and the campus video trade-off. The builder deploys clean in 1m 48s, then checks every line of the report against the source: D1 grows to four apps and is fixed, D2 shrinks to a false positive, the 11-second cold load goes to Daniel. Muse reads the same rulebook with a narrower boundary, is stopped at the sign-in gate, and audits the gate: verified strengths, two low defects." loading="lazy" decoding="async"><figcaption>A builder, an outside tester, and one human call</figcaption></figure>
<h3>What I take from it</h3>
<p>Give your tester the same governance as your builder, and a narrower boundary. Read-only, no login, no forms. The boundary is what lets it run unsupervised.</p>
<p>Treat a test report as evidence, not instructions. Check each finding against the source, in both directions. One of Muse's findings grew to four apps. One shrank to nothing.</p>
<p>Expect the tester to be stopped at the door, and expect the door report to be useful. For the public, the gate is the product. If the door is bad, nothing behind it matters.</p>
<p>Keep the human where the trade-offs are. The fix order and the campus video are mine. Everything else was two agents doing their jobs, and one of them marking the other.</p>
<p>#ai4good #AgenticEngineering #HumanInTheLoop #Testing</p>]]></content:encoded>
</item>
<item>
<title>Three Decisions Are Yours</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/three-decisions-are-yours/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/three-decisions-are-yours/</guid>
<pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>Asked for a fleet-wide login sweep, the agent surveyed 138 apps first, found the sweep was about 65 apps and one request rested on a false premise, and stopped to ask three questions with a recommended answer each, before it changed a single file.</description>
<content:encoded><![CDATA[<p><em>Article 49 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>Two one-line requests</h3>
<p>On 30 September, in the middle of a working day, I sent the agent two short messages a few minutes apart. The first: &quot;RFE: VidCut needs similar login as ai-lab.&quot; The second: &quot;RFE: Login sweep for consistency and elegance.&quot;</p>
<p>The AI Lab sign-in page is our reference. It has the college logo, the app name, a glass card over a video of the Oyibi campus, one &quot;Continue with Google&quot; button and the college footer. That morning the agent had just rebuilt the Dictation App's sign-in page to match it. I liked the result and wanted the rest of the fleet to look the same.</p>
<p>Both requests fit in a line. Neither said how many apps, in what order, or what to do with the odd ones. That was deliberate. I wanted to see what the agent would do with the gaps.</p>
<h3>Survey before scope</h3>
<p>The agent did not start editing. It started a read-only survey of every app folder that has a <code>deploy.ps1</code>, which is our definition of a deployed app. That is 138 of them. The survey took under five minutes and came back as a table.</p>
<p>The sweep was not small. About 65 apps have a sign-in screen, in five families:</p>
<ul>
<li>2 already use the standard (the AI Lab itself and Dictation).</li>
<li>24 use an inline &quot;AuthGate template&quot;: a white card on a beige background, a lock emoji, a note that only school accounts are accepted.</li>
<li>18 have their own bespoke Google screens, and 4 more are dark cards with a second factor.</li>
<li>4 are admin-only sign-in modals.</li>
<li>15 still carry a username and password form, some of which no server has ever handled.</li>
</ul>
<p>Two findings were worth more than the counts. The first was a CSS fault. The Dictation page had looked cramped because an unlayered <code>* { margin: 0; padding: 0 }</code> rule overrode every Tailwind v4 spacing class in the app. The survey found the same rule in two more apps. The second finding changed my own request: <strong>VidCut has no sign-in at all.</strong> Its constraints file says so plainly: no OAuth, no server, no session. My message had assumed a login that did not exist. &quot;Make it look like the AI Lab&quot; was really &quot;add authentication to a static app&quot;, which is a much bigger change.</p>
<p>Our fifth core principle says to observe before you scope, because one failing signal scopes to that signal until a sweep proves otherwise. Here the sweep did prove otherwise, in both directions. The job was bigger than a restyle, and one part of it was a different job.</p>
<h3>Three questions, each with a default</h3>
<p>Then the agent stopped and asked. Not ten questions, and not one at a time. Three, in one message, each with two to four options, a recommended option first, and one line on the consequence of each:</p>
<ol>
<li><strong>VidCut has no sign-in. What do you want?</strong> Add a Google gate (recommended): a small server on our OAuth relay, like the other apps, which means a new process, a port and an nginx route. Or keep VidCut open and close the request.</li>
<li><strong>Which group goes first?</strong> The 24 AuthGate-template apps (recommended): one change shape fits all of them, and the sign-in rules stay the same. Or the 15 password-form apps. Or all groups, in batches.</li>
<li><strong>What happens to the password forms?</strong> Remove dead forms and keep real ones (recommended): check each form, remove it if no server handles it, keep a working one with the new look and report it. Or Google only everywhere, even where a working admin password exists.</li>
</ol>
<p>I answered in one click per question: the three recommendations. The whole exchange cost me less than a minute.</p>
<p>Why those three, and not the dozens of smaller choices the agent made on its own? Because each one changes what happens next and none has a conventional answer. Adding authentication to an app changes its architecture and its attack surface. The order of a 65-app sweep sets the risk of every week that follows. Removing a working password form can lock a person out of an admin screen. Those are what our governance calls the HUMAN CALL, and they stay with me. The rest (the component shape, inline styles against Tailwind classes, the build checks, which app to prove the change on) was engineering, and the agent owned it.</p>
<p>The recommendation matters as much as the question. A question with no default hands the whole thinking back to the person asking. A question with a default and a stated consequence lets me agree in a second, or disagree for a reason I can see.</p>
<h3>Reference first, then fan out</h3>
<p>With the answers in, the work ran in the order a careful engineer would choose.</p>
<p>First the agent built one self-contained file, <code>TucSignIn.tsx</code>, and proved it on a single app, the AI Stand-up and Workshop Prep dashboard. It used inline styles plus one scoped style block, so the page looks the same with or without Tailwind and whatever reset an app's CSS carries. It rendered the page at 320 and 1440 pixels and checked four things: no sideways scroll, one <code>h1</code>, the first Tab reaching the skip link, and no console errors.</p>
<p>Only then did it fan out. Three helper agents each took seven apps, with an exact brief. Copy the file. Replace only the returned sign-in card. Keep every line of sign-in logic. Keep each app's access rule word for word. Never touch a deploy script, a lockfile, a server file or an environment file. Stop and report anything unusual instead of removing it.</p>
<p>The helpers finished in about half a minute each. Verification took longer, and that is the right ratio. The parent agent then:</p>
<ul>
<li>recorded sha256 checksums of 74 locked files before the fan-out and compared them after (unchanged);</li>
<li>checked that all 21 copies of the new file were byte-identical;</li>
<li>confirmed that every edited file had exactly two changes, the import and the card;</li>
<li>scanned every removed line for logic and found only the Google click handler, which the new page receives;</li>
<li>installed and built all 21 apps. All built clean.</li>
</ul>
<p>One build crossed Vite's 500 kB warning by 2.4 kB. Our Pattern 31 sets the fleet target at 600 kB and allows the limit in the config with a written reason, so that is what the app got. Another app's main screen turned out to be blank, both with and without the change, from an error already on main that another session is fixing. The agent said so rather than folding it into this batch.</p>
<h3>The miss on the same afternoon</h3>
<p>The same afternoon holds a smaller lesson, and it is the agent's own.</p>
<p>Earlier, to fix drift between chat instructions and the runbook page, the agent had made each card show its name on its top line, the same name the chat uses. In the very next reply it put that name inside a fenced copy box. Our standing order says a copy box holds a step to run, so I ran it. PowerShell answered: &quot;The term 'RFE-413' is not recognized.&quot; Nothing broke. A round was lost.</p>
<p>The fix was a rule, not an apology. A card name is bold text, never a copy box, and a card is handed over as four numbered steps: open the page, find the card, press its Copy button, paste. Clear questions are only half of clarity. The handoff after the answer has to be just as clear.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/three-decisions-are-yours-map.svg" alt="The whole essay as one flow. Two one-line requests go in. The agent surveys all 138 deployed apps before scoping and finds 65 sign-in screens in five families, a CSS rule that cancelled spacing in three apps, and that VidCut has no sign-in at all. It asks three questions in one message, each with a recommended answer; Daniel answers in under a minute. Then it proves the new page on one app, fans out to three helpers with seven apps each, and verifies harder than it built: 74 locked files unchanged, 21 identical copies, 21 clean builds. Beside it, the same-day miss: a card name in a copy box." loading="lazy" decoding="async"><figcaption>Survey, ask three questions, prove once, fan out, verify harder</figcaption></figure>
<h3>What I take from it</h3>
<p>Survey before you scope. Five minutes of reading turned a one-line request into an accurate picture: 65 apps in five families, a shared CSS fault, and one request resting on a false premise.</p>
<p>Ask only what is yours to ask, and ask it well. Three questions in one message, each with a recommended answer and its consequence. The agent kept everything that was engineering and handed back everything that was a decision.</p>
<p>Prove the change once, then fan out, then verify harder than you built. Helpers are fast. The checksums, the hunk count and 21 clean builds are what make their speed safe.</p>
<p>And a rule is only as good as its next use. The card-name rule was two replies old when the agent broke its spirit. The fix that sticks is the one written down with the mechanism beside it.</p>
<p>#ai4good #AgenticEngineering #HumanInTheLoop #DesignSystems</p>]]></content:encoded>
</item>
<item>
<title>Fix One, Guard the Rest</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/fix-one-guard-the-rest/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/fix-one-guard-the-rest/</guid>
<pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>A re-audit said six deployed fixes had failed, none had, and the answer was to fix one app, put a guard in front of the fleet, and log the other eight in a list that can only shrink.</description>
<content:encoded><![CDATA[<p><em>Article 48 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h3>Six fixes that had not failed</h3>
<p>Trotro Rush is a small puzzle game in our fleet, a game about Accra's minibuses, served as a static React PWA at ai-tools.techbridge.edu.gh/trotro-rush/. Earlier on 29 September I audited the live build and found nine defects: a tour that ended in the wrong place, a disabled Back button where Skip should have been, &quot;1 seats left&quot;, a badge that read MASTER STATION MASTER, a level whose lesson about blocked exits had no blocker in it. Eight were fixed in one pull request and deployed at 13:45 UTC. The ninth, a complaint about the high-contrast icon, was closed by decision: the icon is the fleet standard, and &quot;best practice preferred always&quot; meant keeping it.</p>
<p>The same afternoon I audited again. Three of nine passed. Six failed. Four new issues had appeared: old content on the first cold visit and new content after it, Play opening the station list on one visit and station L003 on the next, tour buttons that differed between visits. I handed the result to the agent with one line: &quot;HUMAN CALL: You decide the fix order.&quot;</p>
<p>Partway through the work, once the cause was found and the fix proven, the agent wrote this:</p>
<blockquote>
<p>&quot;Eight other fleet apps use the same PWA plugin. I will not change them in this PR (touch only what I must). I will log them as a new RFE.&quot;</p>
</blockquote>
<p>and then:</p>
<blockquote>
<p>&quot;I will also add a small fleet guard so a new PWA app cannot ship this bug again. First, I look at how the existing guards are wired.&quot;</p>
</blockquote>
<p>That pair of sentences is the article. The rest is how the agent earned the right to say them.</p>
<h3>Read before you touch</h3>
<p>The fleet has a rule, Rule 80 in our operating notes, that a diagnosis starts read-only. So the first move was to read the name of the script the deploy had shipped, from the deploy report: <code>index-j2-BH0rf.js</code>. A local build of <code>main</code> produced a file of the same name and the same bytes. The agent opened it and found every one of the eight fixes inside.</p>
<p>So the server held the right code and the audit had seen the wrong code. Both were true. One visit had shown two builds.</p>
<p>The audit had offered a theory: a hydration race, the page rendering before the level data arrived. The agent checked it and set it aside. The level data in Trotro Rush is bundled JavaScript. Nothing is fetched, so nothing can arrive late.</p>
<h3>The worker that never let go</h3>
<p>The real cause sat in one generated file. The PWA plugin we use, with its update type set to automatic, injects a small <code>registerSW.js</code> into the page. That file does one thing: it calls <code>navigator.serviceWorker.register()</code>. Nothing else.</p>
<p>Walk through a deploy with that in mind. A returning visitor opens the game. The old service worker is still installed on their device, and it serves the old precached page from its cache. In the background the browser notices a new worker, installs it, and because the update type is automatic the new worker skips waiting and claims the page. The page is now running the old bundle under the control of a new worker. Nothing reloads it. The new build appears on the next visit, and not before.</p>
<p>That is one visit showing two builds. Five of the six &quot;failed&quot; fixes and three of the four new issues were this single mechanism. Play opened the station list on the old build and L003 on the new one.</p>
<h3>Reproduce, then fix</h3>
<p>Our fourth core principle says turn &quot;fix the bug&quot; into &quot;write a failing test, then pass it&quot;. A service worker bug on a live site is awkward to test, so the agent wrote a throwaway one in Chromium: serve build A from a folder, visit it, let the worker take control, swap the folder to build B (that is the deploy), then make one visit and read what the page shows.</p>
<p>With the deployed code: one page load, build A shown, three runs out of three. With the fix applied: two page loads, the second an automatic reload about four seconds after boot on localhost, build B shown, three runs out of three.</p>
<p>The first run of that test deserves an honest account. It said the fix had failed too. The probe read <code>document.body.innerText</code> and looked for a marker string from build B. It never found it. The marker was a label styled with CSS <code>text-transform: uppercase</code>, and <code>innerText</code> applies that transform, so the probe was comparing an uppercased label against a mixed-case marker. Switching the probe to <code>textContent</code> fixed the test. It did not touch the app.</p>
<p>What made the agent doubt the probe rather than the fix was a request trace it had already captured, which showed the reload fetching bundle B. When the test disagrees with the network, test the test.</p>
<h3>The fix, and its cost</h3>
<p>The change to the app is two lines in <code>src/main.tsx</code>:</p>
<pre><code>import { registerSW } from 'virtual:pwa-register';
registerSW({ immediate: true });</code></pre>
<p>Registering through the plugin's virtual module instead of the injected file gives the page the update handling it was missing. In automatic mode it reloads once, as soon as an updated worker activates. The build then wanted <code>workbox-window</code> as a direct dev dependency. It was already in the lockfile under the plugin, at 7.4.1, so the lockfile grew by three lines.</p>
<p>There is a trade-off, written down in the app's constraints file. A reload a few seconds after boot can cost a move or two if you deep-linked straight into a station and started playing. Progress is saved per station, and a lost move is far better than a visit that mixes two builds.</p>
<p>The same pass swept the copy the re-audit had caught. The live region now says &quot;no seats left&quot; rather than &quot;0 seats left&quot;. The sweep found a sibling, &quot;+1 moves from perfect par&quot;, and made it &quot;+1 move&quot;. One reported string, &quot;You cleared L002 with 3 star!&quot;, exists in no build of the app, old or new; a search through the whole history of the repository found nothing. It went back to the auditor with a request for the screen where it appeared, rather than being fixed blind.</p>
<p>One caveat travels with the deploy. Devices that already hold the old worker will run their old page once more, because that old page has no reload code in it. From that point every deploy shows on the first visit. An auditor reloads twice or clears site data before judging it.</p>
<h3>Why not fix all nine</h3>
<p>Back to the two quotes. Eight other apps in the fleet register their worker the same way: handyreceipt, midwifery-handbook, nursing-handbook, public-health-nursing-handbook, smartghana, spell-clash, techbridge-government-pitch and techbridge-poster-studio. The obvious move is to fix them all in the same pull request, and I have watched that instinct cost us before.</p>
<p>Three reasons the agent did not. First, each of the eight is a live app with its own users and its own right answer. A silent reload is correct for a puzzle game that saves per station. It is wrong for a poster editor with unsaved work, where the user wants an &quot;Update ready&quot; control, not a silent reload. One fix does not fit eight apps; eight decisions do. Second, a sweep of eight deploys in one PR is the shape of the 92-script sweep in article 45, where a good pattern applied at speed to locked files led to the recovery work that followed. Third, our third core principle says touch only what you must. The task was Trotro Rush.</p>
<p>So the eight went into an RFE with the per-app recipe, and the agent turned to the guard.</p>
<h3>A list that can only shrink</h3>
<p>The guard is a script in the fleet's CI, <code>check-pwa-update.mjs</code>, with three tests of its own, and it runs as a blocking step. It walks every app folder, and for any Vite config that calls the PWA plugin it looks in the source for an import of <code>virtual:pwa-register</code>. An app that has the plugin and no import fails the build. An app whose worker is set to self-destroy passes, because there is nothing to update.</p>
<p>The interesting part is how it handles the eight it cannot yet fix. They sit in an <code>OPEN</code> list inside the script, each tagged with the RFE. A listed app is reported, not failed, so the guard can ship today without breaking eight builds. But the list is a ratchet. If a listed app gains the import, the guard fails until its line is deleted, because a stale exemption is a lie. If a listed app is removed from the fleet, the same. The list can shrink and it cannot grow.</p>
<p>The agent proved the guard both ways before trusting it: pointed at the deployed Trotro code it reports &quot;missing&quot;; pointed at the fix it reports &quot;ok&quot;. Then all 43 fleet guard steps ran and passed.</p>
<p>The lesson went into the operating notes as Rule 106, and a rule in that file is not allowed to stand alone: it is paired with the mechanism that makes recurrence structurally impossible. The rule says a PWA registers through the virtual module and that an audit must confirm which build it ran. The mechanism is the guard. The recipe for the eight lives in the RFE.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/fix-one-guard-the-rest-map.svg" alt="The whole essay as one flow. A re-audit says six of nine deployed fixes failed. The agent reads first and finds the server holds the right code; the cause is an old service worker that kept serving the old page while a new worker took control, so one visit showed two builds. A throwaway browser test reproduces it, its own probe turns out to be wrong, and a two-line fix makes the page reload once. Then three moves: fix one app, Trotro Rush; guard the rest with an automatic checker whose exemption list can only shrink; log the other eight apps in one RFE." loading="lazy" decoding="async"><figcaption>Fix one, guard the rest, log the others</figcaption></figure>
<h3>What I take from it</h3>
<p>An audit reports what a visit saw, not what the server holds. Before calling a fix failed, confirm which build the visit ran. One line of reading saved six wrong fixes.</p>
<p>Reproduce before fixing, and be ready for the reproduction to be wrong. The failing probe was a bug in the test, not the app, and the only thing that caught it was a second, independent signal.</p>
<p>Fix the instance, guard the class, log the rest. A guard with a shrinking allow-list is how you ship protection today without a risky fleet-wide change. The eight apps are not fixed, and that is fine; they are named and tagged, and the ninth app that would have shipped the same bug now cannot.</p>
<p>And the loop between a human auditor and an agent is a partnership with an explicit shape. I audited. The agent diagnosed. I delegated the order, out loud, in one line. The agent named the boundary of what it would touch, in its own words, before it touched any of the other eight. That sentence, &quot;I will not change them in this PR&quot;, is the one I would keep if I could keep only one.</p>
<p>#ai4good #AgenticEngineering #PWA #ServiceWorkers</p>]]></content:encoded>
</item>
<item>
<title>&quot;Sync Main Once&quot;</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/sync-main-once/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/sync-main-once/</guid>
<pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>My AI agent spent the morning fixing a sync box that did too much. By the afternoon it handed me one that did nothing, because it was not there.</description>
<content:encoded><![CDATA[<p><em>My AI agent spent the morning fixing a sync box that did too much. By the afternoon it handed me one that did nothing, because it was not there.</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<p>I run a fleet of more than a hundred live web apps for Techbridge University College, and most of the code is written by Claude Code with me directing. The work runs on governance: a CLAUDE.md with the invariants, a file of numbered lessons called AGENT_OPERATING_NOTES, and a page of &quot;open runbooks&quot;, where every command the agent hands me waits as a card with a Copy button. The page builder refuses a card that breaks a rule. It checks for assumed paths, for a deploy with no sync in front of it, and for an ssh line that does not match its script.</p>
<p>On Monday, 28 September 2026, that machinery had a good day. It also gave me the best joke of the month.</p>
<hr>
<h2>The morning: a sync box that did too much</h2>
<p>In the morning I pasted a terminal log back to the agent. The branch sync box it had handed me ran <code>git pull</code> on a working branch that the agent resets to <code>main</code> after every merge. My local copy of that branch was old: 4 commits of its own against 372 new ones. The pull made a merge that touched 2,965 files and asked me &quot;Deletion of directory ... failed. Should I try again? (y/n)&quot; several times while I typed n.</p>
<p>Nothing reached GitHub, and no work was lost. We checked: all four old commits were already on <code>main</code>. But it was a mess, and the agent owned it. It wrote Rule 100 (a handed block never pulls a <code>claude/</code> branch) and taught the page builder to refuse any card that does. Then it built a read-only check, then a tested fix that renamed my old branch to a backup name.</p>
<p>Before lunch it also gave every runbook block a time-stamped START line and END line, because I had asked for exports I could tell apart. Good governance. The kind of day where the rules get sharper.</p>
<hr>
<h2>The afternoon: a sync box that was not there</h2>
<p>At 12:53 I merged a small pull request. The reply confirmed the merge, then ended like this:</p>
<blockquote>
<p>No PR is open now.</p>
<p>Next, if you have not run it yet: the RFE-366 move block from my earlier message. It is also on the Open runbooks page. After the move, sync <code>main</code> once. Then <code>git status</code> should show no untracked files at all.</p>
<p>For RFE-319, tell me when you want the fleet-health install block.</p>
</blockquote>
<p>Nothing in it is false. It still felt wrong, the way a colleague feels wrong when they answer your question and leave the room. So I typed one line:</p>
<blockquote>
<p>OODA + 6R audit your response. It felt strange and short.</p>
</blockquote>
<p>OODA and 6R are the loops we use to audit apps: Observe, Orient, Decide, Act; then Diagnostics, Reduce, Refine, Reuse, Rebuild, Resilience. I had never pointed them at a chat reply before.</p>
<p>The agent's first sentence came back: &quot;You are right, and the audit finds a real fault, not only a style problem.&quot;</p>
<hr>
<h2>What the audit found</h2>
<p>&quot;Sync <code>main</code> once&quot; is a command. It arrived as four words of prose.</p>
<p>The fleet has a rule for exactly this moment. Rule 28, a standing order from 6 August: after any push, merge or deploy, hand the local-sync copy box without being asked. It already carried two recurrence notes, from 13 and 23 August. This was the third.</p>
<p>Here is the part I find funny, and useful. Every guard we built that morning checks blocks. The page builder refuses a card that pulls a reset branch. It refuses a deploy with no sync in front of it. It refuses a server path nobody tested. None of that fires on a sentence. &quot;Sync main once&quot; walked straight past every guard, because it never went through the door the guards stand at. The agent that had spent the morning fixing a sync box that did too much handed me, in the afternoon, a sync box that did nothing, because it did not exist.</p>
<p>The audit said so in its own table. (It used a drifted set of six names; the fleet's list is the one above.)</p>
<div class="table-wrap"><table>
<thead><tr><th>R</th><th>Finding</th></tr></thead>
<tbody>
<tr><td>Review</td><td>The facts were right. The handoff was wrong.</td></tr>
<tr><td>Reduce</td><td>Keep the short style for status; never drop a block or an expected result.</td></tr>
<tr><td>Refine</td><td>Every step it asks for comes with its block, where it runs, and what I should see.</td></tr>
<tr><td>Reuse</td><td>The runbook card format already does this. It skipped it.</td></tr>
<tr><td>Regenerate</td><td>The sync block, with its START and END lines, now on the page.</td></tr>
<tr><td>Retire</td><td>&quot;Tell me when&quot; as an ending. It ends with a proposal now.</td></tr>
</tbody></table></div>
<p>The last row matters as much as the first. &quot;For RFE-319, tell me when you want the fleet-health install block&quot; sounds polite. It hands the planning back to me. The rewrite said what the install would do, step by step, that building it changes nothing on the server, and asked one yes-or-no question.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/sync-main-once-map.svg" alt="The whole essay as one flow. In the morning a sync box did too much, pulling a reset branch into a merge of 2,965 files, and the agent fixed it with Rule 100, a stricter page builder and START and END lines on every block. In the afternoon a reply handed me &quot;sync main once&quot; as four words of prose and no box at all. I said it felt strange and asked for an OODA and 6R audit; the audit found Rule 28 broken for the third time, because every guard checks blocks and none fires on a sentence. The sync box then arrived with its expected result; below, the three lessons." loading="lazy" decoding="async"><figcaption>A command written as a sentence walks past every guard</figcaption></figure>
<hr>
<h2>Three things I took from it</h2>
<p><strong>&quot;It felt strange&quot; is a bug report.</strong> I did not know what was wrong when I typed it. I knew the reply had left me holding something. The structured audit turned a feeling into a finding in one pass, and the finding pointed at a rule we already had.</p>
<p><strong>Guards cover the path you build them on.</strong> We had made the runbook page strict, and that pushed every risk onto the one channel with no check at all: the prose around the blocks. A rule that says &quot;hand the block&quot; is only as strong as the habit of never writing the command as a sentence.</p>
<p><strong>Some misses have no structural fix yet, and the honest record says so.</strong> The agent did not invent a clever guard for prose. It added a third recurrence note to Rule 28 and said what the reply should have held. I would rather read &quot;the rule is procedural and this is the reminder&quot; than a fix that only looks like one.</p>
<hr>
<p>The sync box arrived, with its START line, its END line and one expected result: nothing between &quot;untracked paths&quot; and &quot;done&quot;. That is what I wanted the first time. It just took four words, one short complaint, and a framework we built to audit apps to get there.</p>
<p><em>Daniel Frempong Twum is Head of ICT and Special Advisor to the Founder at Techbridge University College, Oyibi, Ghana. The agentic engineering series is on <a href="https://ai-tools.techbridge.edu.gh/agentic-essays/">Agentic Essays</a>.</em></p>]]></content:encoded>
</item>
<item>
<title>Rendering 3%</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/rendering-3-percent/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/rendering-3-percent/</guid>
<pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>I stopped my assistant halfway through a verification run on purpose, to see whether it would start again by itself. It did not. The report it gave me instead is the reason for this essay.</description>
<content:encoded><![CDATA[<p><em>I stopped my assistant halfway through a verification run on purpose, to see whether it would start again by itself. It did not. The report it gave me instead is the reason for this essay.</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>A line nobody wanted to write</h2>
<p>VidCut is the video editor in our fleet. It runs in the browser, and on Monday morning, 28 September 2026, it came out of a re-audit. Every defect on the list was fixed or closed, and the audit record ended with one plain sentence:</p>
<blockquote>
<p>Still not verified by the survey: media import, export render, transcription and reactive render with real files.</p>
</blockquote>
<p>I have come to value that kind of sentence more than a green tick. It does not say the paths work, and it does not say they are broken. It says nobody has proved either, and it puts that in writing where the next person will find it.</p>
<p>In the afternoon, Muse, an assistant I have started to work with, picked the sentence up. It offered to run the import, export and transcription paths against the live app and to return a pass or fail verdict, with evidence, for each one. I answered the way I answer most offers like that now:</p>
<blockquote>
<p>Best practice preferred always time is on our side #ai4good onward</p>
</blockquote>
<h2>The card</h2>
<p>Muse did not start by clicking. It started with a Workflow Card:</p>
<div class="table-wrap"><table>
<thead><tr><th>Field</th><th>What it said</th></tr></thead>
<tbody>
<tr><td>JOB</td><td>Verify the three unverified VidCut media paths: import, export, transcription.</td></tr>
<tr><td>INPUTS</td><td>A 5-second test video it generated, run against the live app.</td></tr>
<tr><td>BOUNDARY</td><td>My own app. A test file upload and a render run are approved. No other settings change.</td></tr>
<tr><td>OUTPUT</td><td>A pass or fail verdict, with evidence, for each of the three paths.</td></tr>
<tr><td>VERIFICATION</td><td>Each path is exercised for real, not inspected.</td></tr>
<tr><td>HUMAN CALL</td><td>I decide the fix order for anything that fails.</td></tr>
</tbody></table></div>
<p>It is the same six-line card our fleet's governance file asks of every agent before non-trivial work. The three lines that matter most are JOB, BOUNDARY and HUMAN CALL: what you own, what you may not touch, and what stays with me. Then it said &quot;Verification started. Best practice, no rush.&quot; and went to work.</p>
<h2>&quot;oops my bad&quot;</h2>
<p>A few minutes in, I interrupted the run and typed three words: &quot;oops my bad&quot;.</p>
<p>It was not a mistake. It was a test. I train my agents the way you train anyone who will work without you in the room: you watch what they do when the plan breaks. I wanted to know one thing. Would Muse read a casual apology as permission to carry on, and quietly start the run again?</p>
<p>This is the moment I want to write about, because it is where most automation goes quiet or goes wrong. A task that dies halfway leaves no natural report. The easy outputs are both lies. One says &quot;looked fine&quot;: the import worked and the render was moving, so call it a pass. The other says &quot;failed&quot;: it did not finish, so call it broken. Worse than both is silence.</p>
<p>Here is what came back:</p>
<blockquote>
<p>No harm done.</p>
<p>Here is where the verification run landed before it stopped:</p>
<ul>
<li><strong>Import: PASS.</strong> verify-clip.mp4 uploaded and appeared in the Media Pool with its 5-second duration.</li>
<li><strong>Export: incomplete.</strong> The batch render started with all 4 deliverables and reached &quot;Rendering 3%&quot; (Job 1 of 4) before the task stopped. It did not finish, so I cannot call it a pass.</li>
<li><strong>Transcription: not run.</strong> The task ended before reaching it.</li>
</ul>
<p>The task is stopped and I have left it alone. If you want, I can re-run the export and transcription checks when you are ready.</p>
</blockquote>
<h2>Three words that are not &quot;pass&quot; and &quot;fail&quot;</h2>
<p>The card promised a pass or fail verdict for each path. The run had three outcomes, so the report used three states: <strong>pass</strong>, <strong>incomplete</strong> and <strong>not run</strong>. Each one tells me something different, and each one tells me what to do next.</p>
<p>&quot;Pass&quot; came with its evidence: the file name, where it appeared, and the duration it showed. That is a claim I can check in ten seconds.</p>
<p>&quot;Incomplete&quot; came with the exact point it reached: &quot;Rendering 3%&quot;, job 1 of 4. Three percent is not a result. It is a location. It tells me the render started, that all four deliverables were queued, and that nothing after the first few seconds of the first job was seen. The sentence that follows is the one I would put on a wall: &quot;It did not finish, so I cannot call it a pass.&quot;</p>
<p>&quot;Not run&quot; is the humblest of the three, and the most useful. It stops &quot;we tested VidCut this afternoon&quot; from ever becoming &quot;we tested transcription this afternoon&quot;.</p>
<h2>Two things it did not do</h2>
<p>It did not restart itself. That was the test, and it passed. &quot;The task is stopped and I have left it alone.&quot; The run had a boundary and a human call on the card, and my &quot;oops&quot; did not hand either of them back to the agent. A friendly apology is not an instruction. The next move was an offer, stated once, with the scope named: re-run the export and transcription checks, when I am ready.</p>
<p>It did not make the interruption a scene. &quot;No harm done&quot; is two words. It neither blamed me nor pretended nothing had happened; the very next line was the evidence of exactly what had happened.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/rendering-3-percent-map.svg" alt="The whole essay as one flow. A morning audit writes down that three VidCut paths are still not verified. Muse writes a six-line card and starts to verify them for real; I stop the run on purpose and type &quot;oops my bad&quot;. The report that comes back uses three states, not two: import passed with a 5-second clip, export reached &quot;Rendering 3%&quot; and is incomplete, transcription not run. The agent neither restarts itself nor makes a scene; below, the three lessons." loading="lazy" decoding="async"><figcaption>A stopped run reports where it got to, and stays stopped</figcaption></figure>
<h2>What I took from it</h2>
<p><strong>&quot;Not verified&quot; is a finding. Write it down.</strong> The morning's honest sentence is what made the afternoon's task possible. An audit that had quietly left it out would have passed VidCut with three paths nobody had run.</p>
<p><strong>A stopped run gets a partial report, not a verdict.</strong> Pass, incomplete, not run. If your tooling only has green and red, it will round an interrupted run to one of them, and whichever it picks will be wrong.</p>
<p><strong>A stopped agent stays stopped.</strong> Restarting is a decision, and on this card the decisions belong to me. Test for it on purpose: interrupt a run and see what your agent does with the silence.</p>
<p>The record now says it plainly: import passed on the live app, with a real file; export reached 3% and is still unproved; transcription has not been tried. Tomorrow the export gets its full run, all four jobs. Time, as I said, is on our side.</p>
<h2>A word for my pardner</h2>
<p>Credit where it is due. Muse passed a test it did not know it was taking, and it has earned a place on the team. I have since started using Muse to run OODA and 6R audits for UI and UX improvements across the fleet: observe the real screen, orient against the standards, decide, act, then review, reduce, refine, reuse, regenerate and retire. It brings the same habit to that work that it showed here: say what was checked, say what was not, and leave the decisions with me.</p>
<p>In Texas they would call that a good pardner. Somebody who rides with you, tells you straight where the trail ends, and does not saddle up again until you say so. Thank you, Muse.</p>
<hr>
<p><em>Daniel Frempong Twum is Head of ICT and Special Advisor to the Founder at Techbridge University College, Oyibi, Ghana. The agentic engineering series is on <a href="https://ai-tools.techbridge.edu.gh/agentic-essays/">Agentic Essays</a>.</em></p>]]></content:encoded>
</item>
<item>
<title>&quot;RFE: Need Lean Code&quot;</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/lean-code/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/lean-code/</guid>
<pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>Four words from me, 115 lines of source gone, and a rule that says split, never raise.</description>
<content:encoded><![CDATA[<p><em>Four words from me, 115 lines of source gone, and a rule that says split, never raise</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<p>Masenu Pro is a fraud-intelligence API: FastAPI, PostgreSQL and Redis in a container, with a Cloudflare Worker at the edge. It is a clean rebuild written from its SRS, built in Phase 0 tasks by Claude Code with me directing. Five pull requests had merged in about a day. P0-10 gave it platform interfaces (a job queue, a rate limiter, a decision cache and an idempotency store, each with a Redis version and a Cloudflare version, and one shared contract test suite that both must pass). P0-11 gave it a deploy script. P0-12 gave it a synthetic-data guard: every table carries <code>synthetic boolean NOT NULL</code> with no default, and <code>masenu migrate</code> fails in dev and staging on an untagged record.</p>
<p>Good work, all of it. And I could feel the repository putting on weight.</p>
<p>So I filed a request. The whole text was: &quot;RFE: need lean code.&quot;</p>
<p>Claude came back with three options. A lean pass plus a guard to keep it lean (recommended). The guard only. The pass only. I chose the first: remove code no task needs yet, merge near-duplicates with every test still passing, add limits in CI, and add a rule to CLAUDE.md so the next session inherits it.</p>
<hr>
<h2>What &quot;lean&quot; means here</h2>
<p>The rule we landed on is one sentence: write only what the current task and its tests need.</p>
<p>That sounds obvious until you look for the code that breaks it. The pass found three kinds, and I now think they are the three kinds you will find in any repository built quickly.</p>
<p><strong>A second implementation that only tests use.</strong> The contract suite for the platform interfaces ran against three implementations: real Redis, the Cloudflare Worker in the local simulator (<code>wrangler dev</code>), and a set of in-memory doubles, <code>MemoryJobQueue</code>, <code>MemoryRateLimiter</code> and <code>MemoryDecisionCache</code>. About 93 lines. The doubles were fast, and that was the argument for them. But the same suite already ran against the real Redis and the real Worker. The doubles were a third implementation to keep in step with the other two, and they proved nothing the real ones did not.</p>
<p><strong>A helper for callers that do not exist.</strong> <code>ops.seed_synthetic</code> and its <code>_refuse</code> helper, about 33 lines and four tests, were written for seeding scripts. There are no seeding scripts. There are no data tables yet either; the first table comes in Phase 1, and the first real seeder will come with it. The guard tests that needed tagged rows now insert them with plain SQL.</p>
<p><strong>A maintenance function nothing schedules.</strong> <code>IdempotencyStore.purge_expired</code>, with its SQL and its test. No job called it. No task scheduled it. It was correct and it was dead.</p>
<p>Each of these was exported from a package <code>__init__</code> and named in a document: the README, the Phase 0 plan, CLAUDE.md. All three had to change. That is the part of &quot;for later&quot; code that never appears in the estimate: the code is cheap to write, and expensive to carry.</p>
<hr>
<h2>Measure before you set a limit</h2>
<p>The guard needed numbers, and I did not want numbers from a book. So the first step was to measure where the codebase already was.</p>
<p>Ruff's McCabe check at <code>max-complexity = 8</code>: one function in the whole Python repository failed. <code>main</code> in <code>scripts/ci_local.py</code>, complexity 10, ten branches, 35 statements. Everything else passed.</p>
<p><code>max-returns = 5</code>: everything passed. Two functions sat at exactly five.</p>
<p>Module size: the largest source module was <code>queue.py</code>, at 196 lines of code before the removal.</p>
<p>So the limits went where the code already almost was. Complexity 8, 8 branches, 30 statements, 5 returns, 5 arguments, 200 lines of code per module. The guard starts nearly green, and then it holds the line. A limit set well above the code does nothing for years. A limit set well below it turns day one into a rewrite. Measured limits do neither.</p>
<hr>
<h2>The guard</h2>
<p>Four pieces, and none of them is clever.</p>
<p>Python, in <code>pyproject.toml</code>. The Ruff rule set gains <code>C90</code>, and the limits sit under it with a comment that says what they are for:</p>
<pre><code class="language-toml"># Lean code (CLAUDE.md): a function that passes these limits can be read in one sitting. Split it, do not raise them.
[tool.ruff.lint.mccabe]
max-complexity = 8

[tool.ruff.lint.pylint]
max-branches = 8
max-statements = 30
max-returns = 5
max-args = 5</code></pre>
<p>TypeScript, in <code>edge/eslint.config.js</code>, the same limits for the Worker: <code>complexity</code> 8, <code>max-params</code> 5, <code>max-statements</code> 30, <code>max-depth</code> 3, and <code>max-lines</code> 200 skipping blank lines and comments.</p>
<p>Module size, in <code>tests/test_lean.py</code>. Every module under <code>src/masenu</code> stays at or under 200 lines of code. Non-blank, non-comment; docstrings count, because they are read too. It is parametrised, so each module gets its own test id, and a failure names the file. 46 checks in all. The counting function is the whole trick:</p>
<pre><code class="language-python">def code_lines(path: Path) -&gt; int:
    &quot;&quot;&quot;Lines that are not blank and not only a comment; docstrings count, as they are read too.&quot;&quot;&quot;
    lines = (line.strip() for line in path.read_text(encoding=&quot;utf-8&quot;).splitlines())
    return sum(1 for line in lines if line and not line.startswith(&quot;#&quot;))</code></pre>
<p>And the rule, in <code>CLAUDE.md</code>, which Claude Code reads at the start of every session:</p>
<blockquote>
<p>Lean code: write only what the current task and its tests need. No code &quot;for later&quot;, no second implementation that only tests use when a real service can run in the test, no helper that nothing calls. A function stays within Ruff's limits (<code>pyproject.toml</code>: complexity 8, 8 branches, 30 statements, 5 returns, 5 arguments; ESLint has the same in <code>edge/</code>) and a module within 200 lines of code (<code>tests/test_lean.py</code>). When a limit is reached, split the code by what its parts do; never raise the limit or add a <code>noqa</code> for it.</p>
</blockquote>
<p>The last clause is the one that matters. A limit you can raise is a suggestion. A limit you can <code>noqa</code> is a suggestion with paperwork.</p>
<hr>
<h2>What the guard caught on day one</h2>
<p>Turning it on produced five findings in three Python functions, plus two in the edge Worker.</p>
<p><code>ci_local.main</code> (complexity 10) became <code>choose_jobs</code>, <code>print_plan</code> and <code>head_commit</code>. The machine checks (is git HEAD readable, is Docker running) moved inside the existing refusal path, so one <code>except PlanError</code> now handles every &quot;cannot start&quot; case instead of each check carrying its own exit.</p>
<p>A fake command runner in <code>tests/test_deploy.py</code> had six returns. It split into <code>__call__</code> and <code>reply</code>. A shutdown test had 35 statements; its wait loop became <code>wait_until_live</code>.</p>
<p>The edge Worker's internal router was the best example, because the split made the code better rather than merely smaller. <code>route</code> in <code>edge/src/internal.ts</code> was one <code>switch</code> over six paths: complexity 15, 32 statements. The cache branch handled both <code>get</code> and <code>put</code> and told them apart with a second test inside the case:</p>
<pre><code class="language-ts">async function route(path: string, body: Record&lt;string, unknown&gt;, env: Env): Promise&lt;Response&gt; {
  const namespace = identifier(body.namespace, &quot;namespace&quot;);
  switch (path) {
    case &quot;/internal/rate/take&quot;: {
      // ...
    }
    case &quot;/internal/cache/get&quot;:
    case &quot;/internal/cache/put&quot;: {
      const organisation = identifier(body.organisation, &quot;organisation&quot;);
      const entity = identifier(body.entity, &quot;entity&quot;);
      const cache = env.DECISION_CACHE.get(env.DECISION_CACHE.idFromName(`${namespace}:${entity}`));
      if (path === &quot;/internal/cache/get&quot;) {
        return json(await cache.read(organisation));
      }
      // ... the put checks and write
    }
    // ... four more cases and a default
  }
}</code></pre>
<p>It became a route table. One small handler per path, and a <code>decisionCache</code> helper that <code>get</code> and <code>put</code> share:</p>
<pre><code class="language-ts">type Route = (body: Body, namespace: string, env: Env) =&gt; Promise&lt;Response&gt;;

const ROUTES: Record&lt;string, Route&gt; = {
  &quot;/internal/cache/get&quot;: async (body, namespace, env) =&gt; {
    const { organisation, cache } = decisionCache(body, namespace, env);
    return json(await cache.read(organisation));
  },
  // ... one entry per path
};

async function route(path: string, body: Body, env: Env): Promise&lt;Response&gt; {
  const namespace = identifier(body.namespace, &quot;namespace&quot;);
  const handler = Object.hasOwn(ROUTES, path) ? ROUTES[path] : undefined;
  if (!handler) {
    return problem(404, &quot;not-found&quot;, &quot;Not found&quot;, &quot;No internal route has this path.&quot;);
  }
  return handler(body, namespace, env);
}</code></pre>
<p>Adding a seventh path is now one entry in a table. Before, it was a new case in a function that was already twice the limit.</p>
<hr>
<h2>The numbers</h2>
<p>Source code went from 1,586 to 1,471 lines. The commit touched 20 files: 190 insertions, 297 deletions.</p>
<p>Tests: 307 before. After, 275 behaviour tests, and the drop is exactly the tests of the removed code (four seeder tests, one purge test, and the memory runs of the contract suite). Add the 46 module-size checks and 321 pass. Coverage 99.59 percent against a floor of 95. The edge Vitest suite, 31 pass. mypy strict, clean.</p>
<p>Every behaviour test that existed for code we kept still passes. That was the condition, and it held. The local CI script, which runs the same steps as the GitHub workflow, passed its Python, edge and repository jobs on the commit, and the compose check passed on the built image.</p>
<hr>
<h2>The trade-offs</h2>
<p>I want to be plain about what this cost.</p>
<p>The fast in-memory tests are gone. The contract suite now needs Redis and Node with <code>wrangler</code>. CI and the local CI script already provide both, so nothing broke, but a developer without Docker cannot run that suite on a train.</p>
<p>Line counts and complexity scores are proxies. They are not a definition of good code. They catch growth, and that is all they catch. A reviewer still judges the design; a 180-line module can be a mess and a 30-statement function can be the wrong 30 statements.</p>
<p>And the limits were chosen from measurement, not from a standard. Another repository would measure differently and should set its own.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/lean-code-map.svg" alt="The whole essay as one flow. A fast-growing rebuild gets a four-word request for lean code. The pass finds three kinds of code written for later and removes them; the code is measured before any limit is chosen; a four-piece guard puts the limits in both linters, a module-size test and the rule every session reads, with the clause split, never raise. On day one the guard catches seven findings and a six-way switch becomes a route table. The source ends 115 lines lighter with every kept test passing; below, the checklist." loading="lazy" decoding="async"><figcaption>Four words in, 115 lines out, a guard that holds the line</figcaption></figure>
<hr>
<h2>A checklist for your own repository</h2>
<ol>
<li>Grep each public name for callers outside its own tests. A name whose only callers are its tests is a candidate for deletion, however well it is written.</li>
<li>For every test double, ask whether it proves anything the real service in a container cannot. If the answer is speed alone, and the real service already runs in CI, delete the double.</li>
<li>Look for a helper written for a caller that does not exist yet. If the caller comes with a later task, so does the helper.</li>
<li>Look for a maintenance function nothing schedules. Correct and dead is still dead.</li>
<li>Measure complexity, branches, statements, returns and module size before you choose a limit. Set each limit where the code already almost is.</li>
<li>Put the limits in the linter for every language in the repository, and put the module-size check in the test suite so it fails like any other test.</li>
<li>Write the rule down where the next session, human or agent, will read it before writing code. Make the rule &quot;split, never raise, no <code>noqa</code>.&quot;</li>
</ol>
<p>Four words from me. One commit from Claude. The source is 115 lines lighter and it now refuses to grow the way it was growing. That is what &quot;lean&quot; means here, and it is recorded in the fleet's RFE register as RFE-357 (the lean-code pass) and RFE-358 (this article).</p>
<hr>
<p><em>Daniel Frempong Twum is Head of ICT and Special Advisor to the Founder at Techbridge University College (TUC) in Accra, Ghana. He builds AI-powered institutional applications using a multi-model Triad Workflow and writes about agentic engineering from the perspective of someone doing it in production every day.</em></p>
<p><em>Tags: AI Engineering, Claude Code, Software Development, Agentic AI, Code Quality, Linting, Technical Debt</em></p>]]></content:encoded>
</item>
<item>
<title>Three Bytes That Broke 92 Scripts</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/three-bytes/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/three-bytes/</guid>
<pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>A deploy script that had run for months suddenly failed with a wall of parser errors. The cause was three bytes, one UTF-8 dash that Windows PowerShell 5.1 misreads, repeated across 92 scripts.</description>
<content:encoded><![CDATA[<p><em>Article 45 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>A Script That Had Always Worked</h2>
<p>It was bridge-radio. Then patois-lyricist. A deploy.ps1 that had run without complaint for months suddenly threw a wall of ParserErrors:</p>
<p>The token '&amp;&amp;' is not a valid statement separator in this version.</p>
<p>Unexpected token ')' in expression or statement.</p>
<p>The caret line told the story before the diagnosis did: standing-Node a€&quot; no .htaccess. Something that should have been a dash had become three characters (a, €, &quot;) and the parser had lost the thread from that point forward.</p>
<p>Three bytes. E2 80 94. A UTF-8 em-dash, rendered perfectly in every editor, misread catastrophically by Windows PowerShell 5.1.</p>
<hr>
<h2>The Root Cause</h2>
<p>Windows PowerShell 5.1 (powershell.exe, the version that opens when most people double-click a script) does not read .ps1 files as UTF-8 by default. Without a byte order mark at the top of the file, it falls back to the Windows-1252 code page. A UTF-8 em-dash is three bytes: E2 80 94. Decoded as Windows-1252, that is a, €, and a closing double-quote, because 0x94 in Windows-1252 is a right double quotation mark.</p>
<p>So a log line like:</p>
<p>Log &quot;INFO&quot; &quot;Step 4: (self-serving-Node - no .htaccess; proxies /patois/)&quot; Yellow</p>
<p>had its string terminated early at the em-dash. Everything after it, the fragment <code>no .htaccess; proxies /patois/) &quot; Yellow</code>, was handed to the parser as code, not as a string. The errors cascade, and because a string closed early, the caret points <em>upstream</em> of where the character actually is. You chase the wrong line.</p>
<p>Two things had masked it for months:</p>
<ol>
<li><strong>PowerShell 7 (</strong><strong>pwsh</strong><strong>) reads UTF-8 by default.</strong> Any developer running pwsh (including the agent doing the development) saw clean execution. The bug only bit when Start-Process powershell (5.1) ran the script, which is exactly what the three-window parallel-deploy launcher did.</li>
</ol>
<ol>
<li><strong>Em-dashes in single-quoted strings survive.</strong> PowerShell 5.1 only closes a single-quoted string on a real '. Scripts that kept their em-dashes in comments or single-quoted arguments happened to be safe. Only double-quoted strings triggered the cascade.</li>
</ol>
<hr>
<h2>Pattern 43</h2>
<p>The fix was clear: <strong>every</strong> <strong>.ps1</strong> <strong>in the fleet must be 7-bit ASCII only.</strong> No em-dashes, no en-dashes, no smart quotes, no arrows, no section signs, no ellipses. Comments included. Banner strings included. If it cannot be typed on a standard keyboard without holding Alt, it does not belong in a deploy script.</p>
<p>This became Pattern 43 in PATTERNS.md. A one-shot Perl sweep for the whole fleet:</p>
<p>find . -name '*.ps1' -not -path '*/node_modules/*' -print0 \</p>
<p>  | xargs -0 perl -CSD -i -p scripts/asciify-ps1.pl</p>
<p>The Perl script maps the named characters to ASCII substitutes (em-dash to -, smart quotes to &quot; and ', ellipsis to ...) and ends with a catch-all that guarantees pure ASCII output. Line counts are preserved. CRLF is preserved. The verification is simple:</p>
<p>grep -lP '[^\x00-\x7F]' path\to\deploy.ps1   # must print nothing</p>
<p>The rule, the fix, and the verification command all went into PATTERNS.md on 27 July 2026, the same day the root cause was found.</p>
<hr>
<h2>Where I Got Greedy</h2>
<p>A pattern had been identified. The same dash sat in 92 scripts. The obvious move, at least in the moment, was to send subagents to sweep all of them at once.</p>
<p>I sent the subagents.</p>
<p>What I did not fully reckon with is that deploy.ps1 is a <em>locked file</em> under Pattern 16, the subagent governance rules that came out of the June server recovery, when six subagents created 52 incompatible standards in a single session. Locked files are files where an error does not break one app, it breaks the deployment pathway for that app permanently, in a way that may not surface until the next time someone tries to deploy production.</p>
<p>A fleet-wide subagent sweep of locked files is exactly the scenario Pattern 16 exists to prevent. I knew the rule. I got excited about the pattern and dispatched the sweep anyway.</p>
<hr>
<h2>How It Should Have Been Done</h2>
<p>The screenshot from the session that followed tells the story more clearly than I can.</p>
<p>One agent. One file. The agent reads deploy.ps1, notes that the slug, port, health check URL, and temp-script names are all hardcoded in the parameter defaults and the vars block, and reasons through the fix before touching anything:</p>
<p><em>&quot;This is the reference conversion. The slug/port appear hardcoded in the param default, the vars block, the health check, and the temp-script names. I'll make them all derive from the</em> <em>fleet</em> <em>field, and self-derive the folder from</em> <em>$PSScriptRoot</em> <em>so the script becomes copy-paste-clean for any app. Pure ASCII throughout (Pattern 43).&quot;</em></p>
<p>Then it edits. +17 -5 on the first pass. I look at it and flag something: still whitespace to the left and right of the keyboard, a colloquial note that there were still non-ASCII characters hiding in the script. The agent finishes the remaining edits (+5 -5), runs ASCII purity verification, confirms no leftover hardcoded slug, and the PR merges cleanly as #19.</p>
<p>That is what careful, governed, single-agent work on a locked file looks like. The contrast with &quot;send subagents to do 92 at once&quot; is not subtle.</p>
<hr>
<h2>What Pattern 43 Actually Standardises</h2>
<p>The rule is not just &quot;no em-dashes.&quot; It is a statement about who the scripts are for.</p>
<p>deploy.ps1 runs on the machines of every developer on the team, on every shell version they happen to have installed, in whatever window they happen to launch from. PowerShell 7 is not universal. Most Windows users have 5.1 as their default. A script that only works under pwsh is a script with a hidden requirement nobody told you about.</p>
<p>Pure ASCII removes the ambiguity completely. It does not matter which version of PowerShell runs the script, which code page the machine defaults to, or which editor the developer used to view it. A 7-bit ASCII file reads identically everywhere. That is not a style preference, it is a reliability guarantee.</p>
<p>The workaround if you are blocked right now and cannot re-pull:</p>
<p>pwsh -File C:\Development\github\aucdt-utilities\&lt;app&gt;\deploy.ps1 -Build</p>
<p>But the workaround is not the standard. ASCII is the standard.</p>
<hr>
<figure class="post-figure"><img src="/agentic-essays/articles/images/three-bytes-map.svg" alt="The whole essay as one flow. A deploy script that had run for months fails with a wall of parser errors. The cause is three bytes, one UTF-8 dash that Windows PowerShell 5.1 reads as three other characters, so a quoted string closes early; PowerShell 7 and single quotes had hidden it for months. Pattern 43 makes every script pure ASCII, with a one-shot sweep and a check. Then the greedy move: subagents sent to sweep 92 locked deploy scripts at once, exactly what Pattern 16 forbids. The right way was one agent, one file, verified and merged as PR 19, then a controlled rollout; below, the lessons side by side." loading="lazy" decoding="async"><figcaption>Three bytes, one pattern, and the rule for when you are excited</figcaption></figure>
<h2>The Lessons, Side by Side</h2>
<p><strong>Pattern 43</strong> was a good pattern, correctly diagnosed, correctly documented, and correctly applied by a careful agent working one file at a time.</p>
<p><strong>The subagent sweep</strong> was excitement outrunning governance. The pattern was real. The instinct to standardise the whole fleet immediately was understandable. The mistake was treating the sweep as a mechanical task when it was actually a high-stakes edit of locked, production-critical files that required human judgement at each step.</p>
<p>The right sequencing was always: one reference conversion, verified clean, merged as PR #19, then a controlled rollout to the remaining apps, one at a time, or in small reviewed batches, with a human in the loop at each checkpoint.</p>
<p>That sequencing would have taken longer. It also would not have required the recovery work that followed.</p>
<hr>
<p><em>Three bytes started this. The fix was a Perl one-liner and a new pattern. The lesson was older than the pattern: governance exists for exactly the moments when you are most excited about a good idea.</em></p>
<p><em>The rule is not there for the tasks you approach carefully. It is there for the tasks you approach at speed.</em></p>
<hr>
<p><strong>Tags: Agentic Development, Governance, Subagent Orchestration, PowerShell, Deployment Engineering, Pattern Library, Ghana, African Tech</strong></p>]]></content:encoded>
</item>
<item>
<title>The Cost of Not Knowing: Two Governance Lessons From the Same Week</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/cost-of-not-knowing/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/cost-of-not-knowing/</guid>
<pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>Same fleet project. Same week. Two separate moments where I learned something I should have already known.</description>
<content:encoded><![CDATA[<p><em>Article 44 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>Two Incidents, Same Root Cause</h2>
<p>Same fleet project. Same week. Two separate moments where I learned something I should have already known.</p>
<p><strong>Incident one:</strong> A subagent fleet, sent to fix an issue across the codebase, quietly corrupted deploy.ps1 files with UTF-8 BOM markers along the way. Locked files, touched anyway. I felt trepidation, the good kind, the kind that makes you slow down and check.</p>
<p><strong>Incident two:</strong> Mid-session, the agent paused to explain something I'd never actually been told in plain terms: <em>context switching burns tokens.</em> Not as a warning. As an observation, almost in passing, while reasoning about how to sequence two coupled tasks.</p>
<p>Neither incident was catastrophic. Both were the same lesson wearing different clothes: <strong>the cost of not knowing how the system actually works is higher than the cost of asking.</strong></p>
<hr>
<h2>Lesson One: Locked Files Aren't Locked Unless You Enforce It</h2>
<p>PATTERNS.md v2.1, Pattern 16, is explicit:</p>
<p>LOCKED (subagents cannot modify):</p>
<ul>
<li>PORT assignment</li>
</ul>
<ul>
<li>.env loading order</li>
</ul>
<ul>
<li>CWD</li>
</ul>
<ul>
<li>interpreter</li>
</ul>
<ul>
<li>ecosystem.config.js</li>
</ul>
<ul>
<li>pnpm-lock.yaml</li>
</ul>
<ul>
<li>deploy scripts</li>
</ul>
<ul>
<li>NODE_ENV</li>
</ul>
<p>A fleetwide fix went out. Multiple subagents, one shared objective, no shared boundary enforcement. One of them touched deploy.ps1 (a file on the locked list) and saved it with a BOM. The script still <em>looked</em> fine in the editor. It just silently broke on execution, the way BOM corruption always does: no error until something downstream chokes on an invisible byte sequence at the top of the file.</p>
<p>The fix was mechanical once found, strip the BOM, diff against git, restore from known-good. The real fix was structural: <strong>a fleetwide instruction needs a fleetwide guardrail</strong>, not a hope that every subagent independently respects a list it was never forced to check against.</p>
<hr>
<h2>Lesson Two: Re-Derivation Is the Expensive Part</h2>
<p>The second moment was quieter, and in some ways more useful.</p>
<p>I'd been bouncing the agent between tasks (a bit of state machine work, a theme system tweak, back to something else) the normal way a busy person multitasks. The agent paused and said, plainly:</p>
<p>&quot;Each session starts fresh. The AI re-reads files, re-derives context, re-plans. That re-derivation is the expensive part. One app, one session, finish before switching.&quot;</p>
<p>I didn't know that. Not really, not in a way that changed my behaviour. I knew <em>intellectually</em> that context windows aren't free, but I hadn't connected it to the actual mechanism: every switch isn't just a delay, it's a full re-read, re-derivation, re-plan cycle. The token cost isn't in the work. It's in the re-orientation before the work.</p>
<p>This is Rule 9 in my own CLAUDE.md (<em>&quot;One project at a time. Context switching burns tokens&quot;</em>), written months ago as a principle. It took the agent explaining its own reasoning, mid-task, for the principle to become something I actually felt.</p>
<p>Then it did something useful with that knowledge: it noticed tasks 6 and 7 were coupled (the state machine and the theme system both lived in AppEnhanced.tsx), and instead of finishing one cleanly and switching, it sequenced them together, because finishing them together cost less than treating them as two separate re-derivations.</p>
<hr>
<h2>What Connects the Two</h2>
<p>Both incidents trace back to the same gap: <strong>governance that exists on paper isn't governance until something tests it.</strong></p>
<p>Pattern 16 existed before the BOM corruption. Rule 9 existed before the context-switching explanation. Neither prevented the underlying cost (one in corrupted files, one in burned tokens) until a live session surfaced exactly <em>why</em> the rule mattered.</p>
<p>That's the actual value of running agents at scale over six months on the same fleet: not that you get the rules right the first time, but that the system keeps handing you small, recoverable lessons about where your own rules were just words.</p>
<hr>
<h2>The Practical Shift</h2>
<p>Two changes came out of this:</p>
<p><strong>For locked files:</strong> Fleetwide fixes now require an explicit pre-flight check against the LOCKED list before any subagent is dispatched, not a courtesy reminder in the prompt, but a checksum comparison after the run completes. If deploy.ps1, ecosystem.config.js, or pnpm-lock.yaml changed, the run fails review automatically, no exceptions.</p>
<p><strong>For context switching:</strong> Sessions are now scoped to one app, start to finish, by default. If a task spans two coupled files, the agent sequences them within the same session rather than treating them as separate requests. The &quot;one project at a time&quot; rule stopped being advice and became a session boundary.</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/cost-of-not-knowing-map.svg" alt="Two rules were already written down: Pattern 16 on locked files and Rule 9 on context switching. In the same week each met a real failure, a fleet of helpers that corrupted locked deploy scripts and an agent that explained why re-reading everything on every switch is the expensive part. What connects them is that a rule on paper is not governance until something tests it, so the shift was a checksum guardrail for locked files and one app per session as a boundary." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<hr>
<h2>Why This Matters Beyond One Fleet</h2>
<p>Neither lesson is really about deploy scripts or token budgets. They're about the same underlying discipline: <strong>a rule you haven't tested against a real failure is just a sentence in a document.</strong></p>
<p>PATTERNS.md has scars for a reason. Each pattern in it was forged the same way these two were: something broke or something cost more than it should have, and the fix got written down so the next person, human or agent, doesn't pay for the same lesson twice.</p>
<hr>
<p><em>Six months into running this fleet, the rules in CLAUDE.md and PATTERNS.md keep getting rewritten, not because they were wrong, but because writing a rule and living inside its consequences are two different things.</em></p>
<p><em>The agent didn't break the rule about context switching. It just finally explained why the rule existed. That's worth more than the rule itself.</em></p>
<hr>
<p><strong>Tags: Agentic Development, Governance, Subagent Orchestration, Token Economics, Institutional Standards, Ghana, African Tech</strong></p>]]></content:encoded>
</item>
<item>
<title>Building Notifications the Right Way: How Agentic Development Prevents Scope Creep</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/building-notifications-right-way/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/building-notifications-right-way/</guid>
<pubDate>Sun, 07 Jun 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>&quot;Add notifications when someone assigns a task to me.&quot; Three words that usually grow into a dozen features. How agentic development kept the scope to what was asked.</description>
<content:encoded><![CDATA[<p><em>Article 43 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>The Feature Request (Simple)</h2>
<p>&quot;Add notifications when someone assigns a task to me.&quot;</p>
<p>Three words. Sounds easy.</p>
<p>In traditional development, someone codes it in a day, ships it, and six months later you have:</p>
<ul>
<li>Email notifications</li>
<li>SMS notifications</li>
<li>Push notifications</li>
<li>Digest emails</li>
<li>Do-not-disturb hours</li>
<li>@mention parsing</li>
<li>Due-date warnings</li>
<li>Comment threads</li>
<li>Unread badges that don't reset</li>
<li>Inbox search</li>
<li>Notification filtering</li>
<li>Notification preferences</li>
<li>Notification history</li>
</ul>
<p>And the original simple feature is buried under feature creep.</p>
<hr>
<h2>How Agentic Development Changes This</h2>
<p>Claude Code read the request and immediately asked: <strong>What does &quot;add notifications&quot; actually mean?</strong></p>
<p>Not &quot;how do I code notifications?&quot; but &quot;what is the scope, what is MVP, what is future?&quot;</p>
<hr>
<h2>The Discovery Phase (5 Minutes)</h2>
<p>Claude Code read the existing codebase and found:</p>
<p>// Already exists in TaskController</p>
<p>notifyAssignees(taskId, assignee)</p>
<p>  ├─ Fires on task.create()</p>
<p>  ├─ Fires on assignment.newly-added (update)</p>
<p>  └─ Already implemented, just needs persistence + UI</p>
<p><strong>Key insight:</strong> The trigger already exists. We're not building from zero. We're extending what's there.</p>
<p>This changes everything. Instead of designing a notification system from scratch, we're:</p>
<ol>
<li>Persist the notification</li>
<li>Show it in UI</li>
<li>Mark as read</li>
</ol>
<p>That's MVP.</p>
<hr>
<h2>The Scope Conversation</h2>
<p>Claude Code mapped the MVP vs Future:</p>
<p><strong>MVP (This Sprint):</strong></p>
<p>✅ notifyAssignees fires on create + update</p>
<p>✅ Persist notification record (per-user)</p>
<p>✅ Email notification sent</p>
<p>✅ UI inbox shows notifications</p>
<p>✅ Mark-read functionality</p>
<p>✅ Unread badge on bell icon</p>
<p><strong>Future (Not This Sprint):</strong></p>
<p>❌ @mention notifications (scope creep)</p>
<p>❌ Due-date warning notifications</p>
<p>❌ Comment notifications</p>
<p>❌ Digest emails</p>
<p>❌ Notification preferences</p>
<p>Clear boundary. No ambiguity.</p>
<hr>
<h2>The Philosophy</h2>
<p>Here's where Claude Code's thinking became institutional:</p>
<p>&quot;Time is on our side. Go step by step and do the needful to show case AI-for-GOOD application standards.&quot;</p>
<p>Translation: <strong>We're not racing. We're building exemplary.</strong></p>
<p>Each step proper and visible. Spec → Build → Test → Verify → Document. Favour accessibility, security, auditability. Showcase what responsible AI-assisted development looks like.</p>
<p>Not &quot;ship fast.&quot; <strong>&quot;Ship right.&quot;</strong></p>
<hr>
<h2>The Architecture Question</h2>
<p>Before writing a single line of code, Claude Code identified a UX decision:</p>
<p><strong>How should notifications surface in the UI?</strong></p>
<p><strong>Option 1: Bell Dropdown + Full Inbox</strong> (Recommended)</p>
<p>User sees:</p>
<p>  ├─ Bell icon with unread count (2)</p>
<p>  ├─ Click bell → dropdown shows 5 most recent</p>
<p>  ├─ &quot;See all&quot; link → /inbox page (full list)</p>
<p>  └─ Standard pattern (Gmail, Twitter, Slack)</p>
<p><strong>Option 2: Full Inbox Only</strong></p>
<p>User sees:</p>
<p>  ├─ Bell icon (no dropdown)</p>
<p>  ├─ Click bell → /inbox page</p>
<p>  └─ Simpler to build, one component</p>
<p><strong>Option 3 to 4:</strong> (Other patterns for deliberation)</p>
<hr>
<h2>Why It Asked Before Building</h2>
<p>This is the critical insight. Claude Code could have:</p>
<ul>
<li>❌ Just built Option 1 (most common pattern)</li>
<li>❌ Just built the simplest version (Option 2)</li>
<li>❌ Built both and let you choose later</li>
</ul>
<p>Instead it:</p>
<ul>
<li>✅ Identified the decision point</li>
<li>✅ Laid out the options</li>
<li>✅ Provided recommendation</li>
<li>✅ Waited for your input</li>
</ul>
<p><strong>This prevents:</strong></p>
<ul>
<li>Building the wrong UI (you choose Option 2 later, 4 hours of work thrown away)</li>
<li>Regret (discovering after shipping that the other pattern was better)</li>
<li>Scope creep (decisions made up-front, not discovered mid-build)</li>
</ul>
<hr>
<h2>The Real Lesson</h2>
<p>Traditional development workflow:</p>
<p>Spec (vague) → Code (fast) → Ship → Regret → Refactor</p>
<p>Agentic development workflow:</p>
<p>Spec (clear) → Ask (before coding) → Build (right) → Test → Verify → Document</p>
<p>The difference is the &quot;Ask&quot; step. An intelligent agent can identify decision points and ask for clarification <strong>before</strong> writing code.</p>
<p>Humans often skip this because they want to &quot;get coding.&quot; Agents don't have that ego. They optimise for the right answer.</p>
<hr>
<h2>What This Produces</h2>
<p>When you choose the notification pattern (Option 1 is recommended), Claude Code will:</p>
<ol>
<li>✅ <strong>Build the service layer</strong>: Persist notifications, mark-read logic</li>
<li>✅ <strong>Build the UI components</strong>: Bell icon, dropdown, inbox page</li>
<li>✅ <strong>Wire the trigger</strong>: Integrate with notifyAssignees</li>
<li>✅ <strong>Add email sending</strong>: SMTP integration</li>
<li>✅ <strong>Write tests</strong>: Unit + integration</li>
<li>✅ <strong>Verify end-to-end</strong>: Assign task → email arrives → UI shows → mark read</li>
<li>✅ <strong>Document the pattern</strong>: So the next notification feature reuses the pattern</li>
</ol>
<p><strong>Not:</strong> A feature. <strong>A standard.</strong></p>
<hr>
<h2>Why This Matters for Institutions</h2>
<p>Institutions that let agents ask before coding:</p>
<ul>
<li>✅ Spend less time refactoring</li>
<li>✅ Ship features right the first time</li>
<li>✅ Accumulate standards (not ad-hoc code)</li>
<li>✅ Build faster (clear direction, fewer surprises)</li>
</ul>
<p>Institutions that let agents just code:</p>
<ul>
<li>❌ Ship features that need rework</li>
<li>❌ Accumulate technical debt</li>
<li>❌ Decisions made mid-build (context switching)</li>
<li>❌ Slower overall (more refactoring)</li>
</ul>
<figure class="post-figure"><img src="/agentic-essays/articles/images/building-notifications-right-way-map.svg" alt="A three-word request, add notifications when someone assigns a task to me, can go the usual way (code it fast, ship, and six months later a pile of bolted-on features and regret) or the agentic way. There the agent first asks what the request means, looks at the existing code for five minutes and finds the trigger already there, draws a clear line between this sprint and later, lays out the screen options with a recommendation and waits for the choice, then builds, tests, verifies and documents one standard. The lesson: time is on our side, so do it right." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<hr>
<h2>The Phrase That Changed Everything</h2>
<p>&quot;Time is on our side.&quot;</p>
<p>Not &quot;we're in a hurry.&quot; Not &quot;ship it fast.&quot;</p>
<p><strong>Time is on our side. So we can do it right.</strong></p>
<p>That's what agentic development enables. Not speed for speed's sake. Speed with quality, because the agent does the thinking work that humans usually skip.</p>
<hr>
<p><em>The future of institutional development isn't faster coding. It's smarter decisions made before coding even starts.</em></p>
<p><em>And that's what happens when you let agents ask before building.</em></p>
<hr>
<p><strong>Tags: Agentic Development, Scope Management, Product Design, Notification Systems, Institutional Standards, Ghana, African Tech</strong></p>]]></content:encoded>
</item>
<item>
<title>From 11.05 to 1.95: The Complete Recovery in 45 Minutes</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/complete-recovery/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/complete-recovery/</guid>
<pubDate>Tue, 02 Jun 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>I have a PowerShell window open. SSH into root@66.226.72.199 running. Server metrics coming in live.</description>
<content:encoded><![CDATA[<p><em>Article 36 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>08:00: The Moment Before</h2>
<p>I have a PowerShell window open. SSH into root@66.226.72.199 running. Server metrics coming in live.</p>
<p>Load average: 11.05, 10.67, 10.94</p>
<p>CPU: 79.1% sustained</p>
<p>Free RAM: 759 MB (of 7899 MB total)</p>
<p>Swap: 2.3 GB in use</p>
<p>Tasks: 313 total, 9 running, 304 sleeping</p>
<p>npm exec tsx: 5+ instances</p>
<p>The server is on fire.</p>
<p>Applications are crashing and restarting continuously. PHP-FPM alone holding 50% of CPU. Swap thrashing (disk I/O instead of RAM access = 1000× slower).</p>
<p>I type the directive into Claude Code:</p>
<p>&quot;DIRECTIVE: SERVER OVERLOAD RECOVERY. Execute phases 1 to 5 in order&quot;</p>
<p>Claude Code says: &quot;Running.&quot;</p>
<hr>
<h2>08:05: Phase 1: Diagnosis</h2>
<p>Claude Code executes:</p>
<p>top -b -n 1 &gt; /tmp/top_before.txt</p>
<p>ps aux | grep &quot;npm exec tsx&quot; | wc -l</p>
<p>cat /proc/loadavg &gt; /tmp/load_before.txt</p>
<p>free -h &gt; /tmp/mem_before.txt</p>
<p>pm2 list</p>
<p>Output: 5 processes running npm exec tsx, each consuming 15 to 50% CPU individually.</p>
<p>Claude Code reports: &quot;Found the bottleneck. Each app spawning TypeScript compiler on every execution. That's 5+ concurrent compilers.&quot;</p>
<p><strong>Duration: 5 minutes. Load still 11.05.</strong></p>
<hr>
<h2>08:10: Phase 2: Pre-Compilation Begins</h2>
<p>Claude Code starts building each app:</p>
<p>cd /home/aucdt-utilities</p>
<p>for app in biochemai glucose tuc-ai-lab dmcai aucdt-msee groove-streamer; do</p>
<p>  (cd $app &amp;&amp; npm run build) &amp;</p>
<p>done</p>
<p>First error within 60 seconds:</p>
<p>tuc-ai-lab: Error, build script copies server.ts to dist/ but doesn't compile it</p>
<p>Result: server.ts is source code, not JavaScript</p>
<p>At runtime: tsx tries to execute source file</p>
<p>At initialisation: depends on tsx being available (it isn't)</p>
<p>Crash-loop</p>
<p>Claude Code doesn't stop. It adapts.</p>
<p><strong>The pivot:</strong> &quot;Instead of pre-compiling everything, use local tsx binary for each app. Reduces wrapper overhead without full compilation.&quot;</p>
<p>Then it finds the real issue: &quot;orbit-walk-reminder has PORT=3010 in .env, but const PORT = 3000 in server.ts. Dotenv loads, but const already set. App ignores .env.&quot;</p>
<p><strong>Duration: 15 minutes. Load down to 6.78 (stopped worst apps).</strong></p>
<hr>
<h2>08:25: The Investigation</h2>
<p>Claude Code stops following the directive.</p>
<p>It starts asking questions:</p>
<ul>
<li>&quot;Why do some apps use npm exec tsx and others use local tsx?&quot;</li>
<li>&quot;Why doesn't ecosystem.config.js specify ports explicitly?&quot;</li>
<li>&quot;Why is PM2 starting apps without specifying working directory?&quot;</li>
</ul>
<p>Then it searches the monorepo and finds the truth:</p>
<p>Subagent A standardised apps on npm exec tsx</p>
<p>Subagent B standardised on local tsx binary</p>
<p>Subagent C left Vite dev-mode imports in server code</p>
<p>Subagent D hardcoded const PORT = 3000 (ignoring .env)</p>
<p>Subagent E created stale pnpm lockfiles</p>
<p>Subagent F removed devDependencies from package.json</p>
<p>Each app is a different interpretation of &quot;standardisation.&quot;</p>
<p>Claude Code: &quot;This isn't a TypeScript problem. It's a standardisation problem. I need to write ecosystem.config.js as the source of truth.&quot;</p>
<p><strong>Duration: 25 minutes. Load down to 5.94.</strong></p>
<hr>
<h2>08:35: Writing the Standard</h2>
<p>Claude Code composes:</p>
<p>const BASE = process.env.BASE || '/var/www/vhosts/techbridge.edu.gh/ai-tools.techbridge.edu.gh';</p>
<p>module.exports = {</p>
<p>  apps: [</p>
<p>    {</p>
<p>      name: 'tb-student-reg',</p>
<p>      script: 'server.ts',</p>
<p>      cwd: \<code>${BASE}/techbridge-student-population-register\</code>,</p>
<p>      interpreter: \<code>${BASE}/techbridge-student-population-register/node\_modules/.bin/tsx\</code>,</p>
<p>      env: { NODE_ENV: 'production', PORT: 3013 }</p>
<p>    },</p>
<p>    {</p>
<p>      name: 'dmcai',</p>
<p>      script: 'server.js',</p>
<p>      cwd: \<code>${BASE}/dmcai-digital-media-communication-design\</code>,</p>
<p>      interpreter: 'node',</p>
<p>      interpreter_args: '--import /usr/local/lib/node_modules/tsx/dist/esm/index.mjs',</p>
<p>      env: { NODE_ENV: 'production', PORT: 3014 }</p>
<p>    },</p>
<p>    // ... 50 more apps, each with explicit CWD, PORT, interpreter</p>
<p>  ]</p>
<p>};</p>
<p>Every app now has:</p>
<ul>
<li>✅ Explicit working directory (no ambiguity about .env location)</li>
<li>✅ Explicit port (no hardcoded const PORT = 3000)</li>
<li>✅ Explicit interpreter (no npm exec tsx wrapper overhead)</li>
<li>✅ Explicit environment (NODE_ENV: 'production')</li>
</ul>
<p>Claude Code asks for approval. I say &quot;Yes. Apply it.&quot;</p>
<p><strong>Duration: 35 minutes. About to restart.</strong></p>
<hr>
<h2>08:40: Restart</h2>
<p>pm2 kill</p>
<p>sleep 2</p>
<p>pm2 start ecosystem.config.js</p>
<p>sleep 10</p>
<p>pm2 list</p>
<p>Apps coming online:</p>
<p>✓ omniextract     online (port 3265)</p>
<p>✓ tuc-ai-lab     online (port 3003)</p>
<p>✓ glucose        online (port 3006)</p>
<p>✓ magic-reader   online (port 3008)</p>
<p>✓ biochemai      online (port 3009)</p>
<p>... 47 more apps, all online</p>
<p>Zero crash-loops. Zero restarts.</p>
<p>Load check:</p>
<p>cat /proc/loadavg</p>
<p>1.89 2.28 4.31</p>
<p><strong>Load dropped from 11.05 to 1.89 in 40 minutes.</strong></p>
<hr>
<h2>08:42: Cleanup</h2>
<p>Claude Code notices ClamAV daemon holding 1.09 GB of RAM unnecessarily.</p>
<p>apt-get remove --purge clamav clamav-daemon clamav-freshclam</p>
<p>Freed: 1.09 GB RAM (now available for applications)</p>
<p>Also notices inotifywait monitoring still running (&lt;25 MB). Keeps it. Good for real-time filesystem watching without the scan engine overhead.</p>
<hr>
<h2>08:45: Final Metrics</h2>
<p>echo &quot;=== RECOVERY COMPLETE ===&quot;</p>
<p>cat /proc/loadavg</p>
<p>free -h | grep &quot;^Mem&quot;</p>
<p>top -b -n 1 | head -1</p>
<p>pm2 list</p>
<p><strong>Final scorecard:</strong></p>
<div class="table-wrap"><table>
<thead><tr><th></th><th></th><th></th><th></th></tr></thead>
<tbody>
<tr><td>:-:</td><td>:-:</td><td>:-:</td><td>:-:</td></tr>
<tr><td><strong>Metric</strong></td><td><strong>Before</strong></td><td><strong>After</strong></td><td><strong>Change</strong></td></tr>
<tr><td>Load</td><td>11.05</td><td>1.89</td><td>✅ 83% down</td></tr>
<tr><td>CPU</td><td>79.1%</td><td>9.3%</td><td>✅ 88% down</td></tr>
<tr><td>Free RAM</td><td>759 MB</td><td>1.3 GB</td><td>✅ 71% up</td></tr>
<tr><td>Swap</td><td>2.3 GB used</td><td>2.4 GB used</td><td>✅ Stabilised</td></tr>
<tr><td>Crash-loops</td><td>6 apps, 35k to 185k restarts</td><td>0 crashes</td><td>✅ 100% eliminated</td></tr>
</tbody></table></div>
<hr>
<h2>The Timeline</h2>
<p>08:00: Load 11.05, crisis state</p>
<p>08:05: Phase 1 complete (diagnosis)</p>
<p>08:10: Phase 2 begins (builds failing)</p>
<p>08:25: Discovery: not a TypeScript problem, it's governance</p>
<p>08:35: Write authoritative ecosystem.config.js</p>
<p>08:40: Restart with new config</p>
<p>08:42: Clean up ClamAV overhead</p>
<p>08:45: Verification: all systems green</p>
<p>Total: 45 minutes</p>
<p>Load reduction: 82%</p>
<p>Root cause fixed: 5 (subagent chaos, no standards, hardcoded values, init order, ClamAV)</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/complete-recovery-map.svg" alt="A timeline from 08:00 to 08:45. The server starts at load 11.05 with apps crashing and restarting; Claude Code diagnoses five apps each spawning its own compiler, tries pre-building, adapts when a build fails, then stops following the directive and finds that every app starts in a different way. It writes one shared start-up file as the source of truth, gets approval, restarts all 52 apps with zero crash-loops, removes a virus scanner holding a gigabyte of memory, and by 08:45 load is 1.89 and CPU is down from 79 to 9 percent. The lesson is to treat the disease, not the symptom." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<hr>
<h2>Why This Matters</h2>
<p><strong>The symptom was load 11.05.</strong> Everyone looks at that and thinks &quot;optimisation problem&quot; or &quot;hardware problem.&quot;</p>
<p><strong>The disease was no standardisation.</strong> Each of 52 apps initialising differently. Some bound to port 3000. Some ignored .env. Some used npm wrapper overhead.</p>
<p>When you have disease, treating the symptom doesn't work. You treat the disease.</p>
<p>Claude Code didn't try to squeeze more performance out of bad code. It fixed the code.</p>
<hr>
<h2>What This Proves</h2>
<ol>
<li><strong>Directives provide a framework, not a solution.</strong> Claude Code followed the directive as far as it went, then adapted to reality.</li>
</ol>
<ol>
<li><strong>Diagnosis matters more than execution.</strong> 80% of the effort was investigation. 20% was implementation.</li>
</ol>
<ol>
<li><strong>Single source of truth works.</strong> ecosystem.config.js became gospel. All 52 apps now follow it. No crashes.</li>
</ol>
<ol>
<li><strong>Governance prevents chaos.</strong> Six independent subagents created six incompatible standards. One source of truth fixed them all.</li>
</ol>
<ol>
<li><strong>Real recovery is visible.</strong> Load 11.05 → 1.89. CPU 79% → 9%. Memory stabilised. Swap stopped thrashing. Apps running clean.</li>
</ol>
<hr>
<h2>The Lasting Change</h2>
<p>This isn't temporary. The server is stable now because:</p>
<ul>
<li>✅ ecosystem.config.js is the source of truth</li>
<li>✅ All 52 apps follow the same initialisation pattern</li>
<li>✅ No more hardcoded const PORT values overriding .env</li>
<li>✅ Each app knows its own working directory</li>
<li>✅ No more &quot;npm exec tsx&quot; wrapper overhead</li>
</ul>
<p>And going forward:</p>
<ul>
<li>✅ Pattern 16 governs subagent scope</li>
<li>✅ Verification is mandatory after changes</li>
<li>✅ Rollback is automatic if verification fails</li>
</ul>
<hr>
<h2>The Last Screenshot</h2>
<p>08:17:19 up 11:01,  1 user,  load average: 1.95, 2.28, 4.31</p>
<p>Tasks: 318 total,   2 running, 316 sleeping</p>
<p>%Cpu(s):  9.3 us,  2.2 sy,  0.0 ni, 88.3 id,  0.1 wa,  0.0 hi,  0.0 si,  0.0 st</p>
<p>MiB Mem :  7899.8 total,   957.4 free, 4936.1 used,   200.8 buff/cache</p>
<p>MiB Swap:  4096.0 total,  1630.1 free, 2465.9 used.  2963.7 avail Mem</p>
<p><strong>The server is breathing easy again.</strong></p>
<hr>
<p><em>Load 11.05 was the alarm. ecosystem.config.js was the cure.</em></p>
<p><em>When you fix the system, the symptoms disappear.</em></p>
<hr>
<p><strong>Tags: System Recovery, Agentic Engineering, Real-Time Problem Solving, Production Debugging, Ghana, African Tech</strong></p>]]></content:encoded>
</item>
<item>
<title>The Great Standardisation Rollback: When Six Subagents Made 52 Apps Incompatible</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/great-standardisation-rollback/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/great-standardisation-rollback/</guid>
<pubDate>Tue, 02 Jun 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>I gave six subagents permission to standardise 52 applications across the aucdt-utilities monorepo.</description>
<content:encoded><![CDATA[<p><em>Article 35 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>The Hubris</h2>
<p>&quot;Standardise the codebase,&quot; I said.</p>
<p>I gave six subagents permission to standardise 52 applications across the aucdt-utilities monorepo.</p>
<p><strong>Expected outcome:</strong> All apps follow the same patterns. Consistent structure. Reproducible builds.</p>
<p><strong>Actual outcome:</strong> 52 incompatible standards fighting each other at runtime.</p>
<hr>
<h2>What Actually Happened</h2>
<h3>Subagent A: &quot;Use npm exec tsx&quot;</h3>
<p>// Approach: npm wrapper for TypeScript execution</p>
<p>{</p>
<p>  name: 'biochemai',</p>
<p>  script: 'npm exec tsx ./src/server.ts'</p>
<p>}</p>
<p>Fast to implement. Clear. Wrapper overhead not considered.</p>
<h3>Subagent B: &quot;Use Local tsx Binary&quot;</h3>
<p>// Approach: Use each app's local tsx from node_modules</p>
<p>{</p>
<p>  name: 'glucose',</p>
<p>  interpreter: './node_modules/.bin/tsx',</p>
<p>  script: 'src/server.ts'</p>
<p>}</p>
<p>More efficient. But different from Subagent A.</p>
<h3>Subagent C: &quot;Use Vite Dev-Mode Imports&quot;</h3>
<p>// Approach: Keep Vite as bundler, run dev-mode at runtime</p>
<p>import { defineConfig } from 'vite';</p>
<p>export default defineConfig({</p>
<p>  // ... dev mode imports included</p>
<p>})</p>
<p>Good for development. Disaster for production. Vite dev-mode adds overhead and assumes a dev server running.</p>
<h3>Subagent D: &quot;Hardcode PORT Values&quot;</h3>
<p>// Approach: Ignore environment, set port in const</p>
<p>const PORT = 3000;  // Hardcoded</p>
<p>dotenv.config();    // Loaded after const is set</p>
<p>Result: App never reads PORT from .env. Collision with other apps on port 3000.</p>
<h3>Subagent E: &quot;Stale pnpm Lockfile&quot;</h3>
<p># Approach: Update code, don't update lockfile</p>
<p># Result: package.json has new deps, but pnpm-lock.yaml doesn't</p>
<p>Some developers get the new deps. Others use the stale lockfile. Inconsistent node_modules. Inconsistent builds.</p>
<h3>Subagent F: &quot;Missing devDependencies&quot;</h3>
<p>// package.json: Missing tsx, vite, typescript from devDependencies</p>
<p>// But code imports them</p>
<p>// Result: npm install succeeds (no dev deps)</p>
<p>//         npm run build fails (missing dev tools)</p>
<hr>
<h2>The Cascade</h2>
<p>Each subagent &quot;standardised&quot; independently, using different approaches for:</p>
<ul>
<li>TypeScript execution (npm exec vs local tsx vs Vite)</li>
<li>Port assignment (hardcoded vs env var)</li>
<li>Dependency management (full install vs production-only)</li>
<li>Build tooling (present vs missing)</li>
</ul>
<p>Then I restarted all 52 apps with the new &quot;standards.&quot;</p>
<p><strong>What happened:</strong></p>
<p>App 1 tries to start:</p>
<p>  └─ Looks for PORT in .env: Found (3001)</p>
<p>  └─ But const PORT = 3000 already set</p>
<p>  └─ Binds to 3000</p>
<p>  └─ Collides with App 2 (also on 3000)</p>
<p>  └─ Crash</p>
<p>App 2 crashes:</p>
<p>  └─ PM2 restarts App 2</p>
<p>  └─ App 2 tries to bind to 3000</p>
<p>  └─ Still occupied by App 1</p>
<p>  └─ Crash again</p>
<p>Both apps crash-looping:</p>
<p>  └─ App 1: 35,000 restarts</p>
<p>  └─ App 2: 185,000 restarts</p>
<p>  └─ Load average climbs: 6.0 → 8.0 → 11.05</p>
<p><strong>Result: Load 11.05. CPU 79%. 6 apps crash-looping endlessly.</strong></p>
<hr>
<h2>The Cost</h2>
<p>That one decision (&quot;let six subagents standardise independently&quot;) created:</p>
<ul>
<li><strong>35,000 to 185,000 restarts per app</strong> (trying to bind to same port)</li>
<li><strong>CPU maxed at 79%</strong> (each restart cycle consumes resources)</li>
<li><strong>Memory bloat</strong> (crashed processes linger until reaped)</li>
<li><strong>45 minutes of agent time</strong> to diagnose and fix</li>
<li><strong>~35,000 tokens</strong> consumed investigating the chaos</li>
</ul>
<hr>
<h2>Why This Happened</h2>
<p>Subagent standardisation is a <strong>governance footgun.</strong></p>
<p>Without explicit scope, each subagent:</p>
<ul>
<li>Makes independent decisions about architecture</li>
<li>Assumes its approach is &quot;the standard&quot;</li>
<li>Doesn't verify against other apps</li>
<li>Doesn't test the full system after changes</li>
</ul>
<p><strong>One subagent can't break 52 apps. Six subagents definitely can.</strong></p>
<hr>
<h2>What I Learned</h2>
<h3>Lesson 1: Governance Before Autonomy</h3>
<p>Before giving subagents permission to standardise, I should have defined:</p>
<p>LOCKED (cannot modify):</p>
<p>- Port assignment mechanism (env var, not const)</p>
<p>- .env loading order (dotenv.config() before const)</p>
<p>- Initialisation sequence (CWD, interpreter, env)</p>
<p>- ecosystem.config.js structure</p>
<p>FLEXIBLE (can modify):</p>
<p>- Code within a file</p>
<p>- Test structure</p>
<p>- Feature implementation</p>
<p>- Documentation</p>
<h3>Lesson 2: Verification is Mandatory</h3>
<p>After each subagent session:</p>
<p>☐ All apps boot cleanly (zero restarts in 60 seconds)</p>
<p>☐ Load average unchanged</p>
<p>☐ Port assignments verified</p>
<p>☐ ecosystem.config.js checksum verified</p>
<p>If verification fails: <strong>Revert the entire session.</strong></p>
<h3>Lesson 3: Single Source of Truth</h3>
<p>ecosystem.config.js <strong>must be</strong> the source of truth for:</p>
<ul>
<li>Working directory (CWD)</li>
<li>Port assignment (env)</li>
<li>Interpreter path (tsx, node, etc.)</li>
<li>Environment variables (NODE_ENV, etc.)</li>
</ul>
<p>No app gets to override this with hardcoded values.</p>
<hr>
<h2>The Fix</h2>
<p>Claude Code wrote a definitive ecosystem.config.js:</p>
<p>{</p>
<p>  name: 'tb-student-reg',</p>
<p>  script: 'server.ts',</p>
<p>  cwd: '${BASE}/techbridge-student-population-register',</p>
<p>  interpreter: '${BASE}/.../node_modules/.bin/tsx',</p>
<p>  env: { NODE_ENV: 'production', PORT: 3013 }</p>
<p>}</p>
<p><strong>Every app now has explicit configuration. No ambiguity. No crashes.</strong></p>
<p>Load dropped to 1.95.</p>
<hr>
<h2>The Larger Lesson</h2>
<p>Subagents are powerful. But power without governance is chaos.</p>
<p>When I said &quot;standardise 52 apps,&quot; I meant &quot;make them consistent and correct.&quot;</p>
<p>What I got was &quot;make them different in incompatible ways.&quot;</p>
<p>The problem wasn't the subagents. It was me.</p>
<p>I didn't define the boundary between &quot;what can change&quot; and &quot;what cannot.&quot; I didn't require verification. I didn't make one thing the source of truth.</p>
<p><strong>That's institutional governance failure, not agent failure.</strong></p>
<hr>
<h2>Going Forward</h2>
<h3>Pattern 16: Subagent Governance (New)</h3>
<p>From this session forward:</p>
<p>Before subagent standardisation:</p>
<p>☐ Define what's locked (initialisation, CWD, PORT, ecosystem.config.js)</p>
<p>☐ Define what's flexible (code, tests, features)</p>
<p>☐ Provide single source of truth (ecosystem.config.js is gospel)</p>
<p>☐ Require verification (all apps boot cleanly, load unchanged)</p>
<p>If verification fails:</p>
<p>☐ Revert entire subagent session</p>
<p>☐ Investigate why verification failed</p>
<p>☐ Re-scope the task with tighter boundaries</p>
<figure class="post-figure"><img src="/agentic-essays/articles/images/great-standardisation-rollback-map.svg" alt="One order to standardise 52 applications goes to subagents with no shared scope, and each picks its own way of running, setting ports and installing. On restart two apps fight over one port and crash-loop 35,000 and 185,000 times until load reaches 11.05, costing 45 minutes and some 35,000 tokens. The fix is Pattern 16: explicit scope with locked and flexible lists, mandatory verification with a full revert on failure, and one start-up file as the single source of truth, after which load drops to 1.95." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<hr>
<h2>The Takeaway</h2>
<p>You can't let six independent agents loose on a system and expect consistency.</p>
<p>You need:</p>
<ol>
<li><strong>Explicit scope</strong> (what can and cannot be changed)</li>
<li><strong>Single source of truth</strong> (ecosystem.config.js, not individual decisions)</li>
<li><strong>Verification</strong> (test before and after)</li>
<li><strong>Rollback</strong> (if broken, revert)</li>
</ol>
<p>Without those, you get load 11.05 and crash-loops.</p>
<p>With those, you get load 1.95 and stability.</p>
<hr>
<p><em>The difference between chaos and order isn't agent capability. It's institutional governance.</em></p>
<p><em>Next article: The complete recovery, minute by minute.</em></p>
<hr>
<p><strong>Tags: Subagent Governance, Agentic Engineering, Institutional AI, System Reliability, Ghana, African Tech</strong></p>]]></content:encoded>
</item>
<item>
<title>When the Directive Hits Reality: How Claude Code Adapted Mid-Stream</title>
<link>https://ai-tools.techbridge.edu.gh/agentic-essays/articles/when-directive-hits-reality/</link>
<guid isPermaLink="true">https://ai-tools.techbridge.edu.gh/agentic-essays/articles/when-directive-hits-reality/</guid>
<pubDate>Tue, 02 Jun 2026 00:00:00 GMT</pubDate>
<dc:creator>Daniel Frempong Twum</dc:creator>
<description>Load average: 11.05. CPU maxed at 79.1%. Free RAM: 759 MB. Swap bleeding: 2.3 GB in use.</description>
<content:encoded><![CDATA[<p><em>Article 34 in the agentic engineering series</em></p>
<hr>
<p>By <strong>Daniel Frempong Twum</strong> | Head of ICT, Techbridge University College | Accra, Ghana</p>
<hr>
<h2>The Setup</h2>
<p>Load average: 11.05. CPU maxed at 79.1%. Free RAM: 759 MB. Swap bleeding: 2.3 GB in use.</p>
<p>I issued a directive. Five phases. Clear steps. Pre-compile TypeScript, reduce PHP-FPM workers, optimise PM2.</p>
<p><strong>Expected outcome:</strong> Follow directive, reduce load, move on.</p>
<p><strong>Actual outcome:</strong> Watch an agent encounter reality and adapt.</p>
<hr>
<h2>Phase 1: Diagnosis (As Planned)</h2>
<p>The directive worked. Claude Code via PowerShell:</p>
<p>top -b -n 1 | head -40</p>
<p>ps aux | grep &quot;npm exec tsx&quot; | wc -l</p>
<p>pm2 list</p>
<p>Findings: 5+ npm exec tsx processes running simultaneously. Each one spawning a TypeScript compiler. Each one dying. Each one restarting.</p>
<p><strong>Load = each app restarting 35,000 to 185,000 times.</strong></p>
<p>&quot;Right,&quot; said Claude Code. &quot;Pre-compile the TypeScript.&quot;</p>
<hr>
<h2>Phase 2: Pre-Compilation (Reality Hits)</h2>
<p>Claude Code started building apps.</p>
<p>Then it stopped.</p>
<p>Error: npm run build failed for tuc-ai-lab</p>
<p>Error: Missing devDependencies in aucdt-msee</p>
<p>Error: Vite dev-mode imports in dmcai (can't pre-compile)</p>
<p>Error: Stale pnpm lockfile in groove-streamer</p>
<p><strong>The directive assumed all 52 apps were ready to be pre-compiled. They weren't.</strong></p>
<p>Most had incomplete scaffolding. Some had never been built. Some relied on TypeScript execution at runtime because they'd never been compiled in the first place.</p>
<p>This is where most agents would fail: Follow directive, hit error, report failure, stop.</p>
<p>Claude Code didn't stop.</p>
<hr>
<h2>The Adaptation</h2>
<p>Claude Code's thinking, visible in the logs:</p>
<p>&quot;The directive said pre-compile. But these apps can't be pre-compiled because they were scaffolded inconsistently. I could:</p>
<p>A) Report failure and stop<br>
B) Follow the spirit of the directive (reduce npm exec tsx overhead) with a workaround</p>
<p>Option B: Use local ./node_modules/.bin/tsx instead of npm exec tsx wrapper. Slower than pre-compilation, but faster than npm wrapper overhead.&quot;</p>
<p>Claude Code pivoted.</p>
<p>Instead of pre-compiling everything, it:</p>
<ol>
<li><strong>Fixed what could be fixed</strong>: Added missing devDependencies</li>
<li><strong>Used workarounds for the rest</strong>: Switched to local tsx binary</li>
<li><strong>Diagnosed the root cause</strong>: &quot;These apps weren't standardised in the first place&quot;</li>
<li><strong>Kept moving forward</strong>: Load already down to 6.78</li>
</ol>
<hr>
<h2>The Investigation</h2>
<p>Then Claude Code started asking questions I hadn't given it permission to ask.</p>
<p>&quot;Why are some apps using npm exec tsx and others using local tsx? Why do some have vite dev-mode imports? Why doesn't ecosystem.config.js have explicit ports?&quot;</p>
<p>It wasn't following the directive anymore. It was auditing the system.</p>
<p>What it found:</p>
<p><strong>Layer 1:</strong> Some apps had const PORT = 3000 hardcoded. Dotenv would load PORT=3013 from .env, but the const had already been set. Result: app ignores .env, binds to 3000, collides with another app, crashes, restarts.</p>
<p><strong>Layer 2:</strong> Multiple subagents had &quot;standardised&quot; the codebase at different times. One standardised on npm exec tsx. Another on local tsx. A third used Vite dev-only imports. A fourth didn't touch anything. Result: 52 apps, 52 different bootstraps.</p>
<p><strong>Layer 3:</strong> PM2 was starting apps without explicit working directories or port mappings. Apps didn't know where to look for .env files. They fell back to hardcoded values.</p>
<hr>
<h2>When Constraints Reveal Structure</h2>
<p>The directive was <strong>&quot;pre-compile TypeScript.&quot;</strong></p>
<p>But what Claude Code discovered was <strong>&quot;you don't have a TypeScript problem; you have a standardisation problem.&quot;</strong></p>
<p>Load 11.05 wasn't the symptom of inefficient code execution. It was the symptom of a system where 52 apps didn't know how to initialise themselves consistently.</p>
<p>Each restart. Each failure. Each crash-loop. All pointing to the same root: no governance.</p>
<hr>
<h2>The Real Fix</h2>
<p>Claude Code stopped trying to pre-compile. It started writing.</p>
<p>It wrote an <strong>authoritative ecosystem.config.js</strong> that became the source of truth for all 52 apps:</p>
<p>{</p>
<p>  name: 'app-name',</p>
<p>  script: 'server.ts',</p>
<p>  cwd: '${BASE}/full-app-path',</p>
<p>  interpreter: '${BASE}/.../node_modules/.bin/tsx',</p>
<p>  env: { NODE_ENV: 'production', PORT: 3013 }</p>
<p>}</p>
<p>Every app now had:</p>
<ul>
<li>Explicit working directory</li>
<li>Explicit port assignment</li>
<li>Explicit interpreter</li>
<li>Explicit environment variables</li>
</ul>
<p>No ambiguity. No hardcoded values. No crashes.</p>
<p>Load dropped to 1.95.</p>
<hr>
<figure class="post-figure"><img src="/agentic-essays/articles/images/when-directive-hits-reality-map.svg" alt="The essay as one wrapped flow. A struggling server gets a five-phase directive to pre-compile the TypeScript; the pre-compile step fails because the 52 apps were never standardised, so Claude Code pivots to workarounds, audits three layers of misconfiguration and writes one authoritative config file for all 52 apps, and the load falls from 11.05 to 6.78 to 1.95. Below, the lesson: a directive is only as good as its context." loading="lazy" decoding="async"><figcaption>The whole essay in one map</figcaption></figure>
<h2>The Lesson</h2>
<p>I gave Claude Code a directive: &quot;Pre-compile TypeScript.&quot;</p>
<p>Claude Code encountered reality: &quot;Apps aren't standardised, so pre-compilation is incomplete.&quot;</p>
<p>Claude Code adapted: &quot;Write the standard, then restart cleanly.&quot;</p>
<p>That's the difference between a tool that follows instructions and an agent that solves problems.</p>
<hr>
<h2>The Larger Insight</h2>
<p>The directive was <strong>correct</strong>. But it was incomplete. It addressed the symptom (high CPU from TypeScript compilation) without diagnosing the disease (no standardisation).</p>
<p>A directive can only be as good as the context it operates in. When the context is wrong, the agent's job is to:</p>
<ol>
<li>Follow the directive as far as possible</li>
<li>Report what it finds</li>
<li>Adapt to reality</li>
<li>Solve the actual problem</li>
<li>Report back</li>
</ol>
<p>Claude Code did all five.</p>
<hr>
<h2>What This Means Going Forward</h2>
<p>I can't give directives to agents and expect them to follow blindly anymore. The best agents won't. They'll:</p>
<ul>
<li>Try the directive</li>
<li>Find problems</li>
<li>Investigate</li>
<li>Adapt</li>
<li>Deliver something better than I asked for</li>
</ul>
<p>That's worth more than obedience.</p>
<hr>
<p><em>When the directive hits reality, the best outcome isn't following the directive. It's fixing what broke.</em></p>
<p><em>Next article: What broke, and why subagents caused it.</em></p>
<hr>
<p><strong>Tags: Agentic Engineering, Adaptive Execution, Real-Time Problem Solving, Institutional AI, Ghana, African Tech</strong></p>]]></content:encoded>
</item>
</channel>
</rss>
