Coding Higher Mind: the AI system behind the work

Two AI tools, hardened prompts, automatic checks at every step, AI reviewers split by speciality, and a loop that learns from failures. Counted, not estimated. Now open source.

PG v1.5.0 (as of October 6, 2026). This is a version that is still maturing, not a finished product: 12 public versions in under four weeks (1.0.0 on September 12, 1.1.0 on the 13th, 1.2.0 on the 27th, six patch releases 1.2.1 to 1.2.6 on the 28th, 1.3.0 and 1.4.0 on October 5, 1.5.0 on October 6). The counters in the table below were taken on September 26 and I have not recounted them since.

My CV says AI coding agents write the code, while I own the spec, the review and the deploy. A claim like that needs evidence, so this page shows the system itself: what runs in the background, what it enforces, how it learns from its own failures, and where its limits are. In short: a change written by AI cannot reach a product without passing automatic checks, the AI has to show proof before it says “done”, and it can use passwords and keys without ever seeing them. Everything below runs today. Every number came from a command run on the day this page shipped: counted, not estimated. The whole system is a public repository you can install on your own machine in five minutes.

github.com/kamiljan11/coding-higher-mind →

2,853agent transcript files
160past failures turned into automatic checks
23hard stops in the git gates
9specialist AI reviewers
52zero-token tools
40routines (8 code · 32 desktop)
33repos under strict protection
99keys and passwords the AI never sees

The whole machine on one map

KAMILspec · architecture · review · deployClaude Coderepos · CI · production appsGitHub · browser · mail via MCPClaude Desktop (Cowork)mail · documents · researchoutreach · design · opssame rules · same memoryVAULT99 secretskeys go to the process,never to the chatMEMORY1,900+ notesloaded at session start,sessions self-documentGATESprompt · edit · command · stop · commit · push · CI · review · production proofMAS GroupWorkshop 3.0Flytkamiljan.comclient sitesI set the direction. Every change passes the automatic checks (gates) before it reaches a product. Proof comes back up.

Mind map: what „higher mind” is made of

The system that watches over every coding session is called PG (Prompt-Guard). “Higher mind” comes from my other long project, a free guidebook on practical spirituality, whose rule is practice over belief. Here that means: a rule you intend to follow is a belief; a check that runs on its own is a practice.

Prompt layerprotocol on every promptcompressed reportsskill routerEvent hooksedit → lint + typescommand guardstop gate: tier from diffGit gatessecrets · base freshnessdeps · duplicates · TODOsSQL lint · boundariesReviewer departmentscode · security · data · opsux · product · qaverifier · devil's advocateDoctrinedesign before code“done” per risk tier160 scarsRepo templateCI + mutation testingparsed boundariesQA matrix · privacy · runbookRoutinesPR reviewer · guard healthCVE watch · calibrationwatchdog · backup · reportsSelf-testsevery gate has a testrule → gate coverageweekly auditHIGHER MINDPG · PROMPT-GUARDEight parts, one rule: nothing important exists only as written text. Every rule has an automatic check.

Three ideas the whole thing rests on

  • Checks, not written rules: the audit that started all this found that every rule written only as text was broken at scale: 6 of 7 repositories had 0 architecture decision records, the changelog was 105 commits behind, code review ran in only 2.5 % of sessions, and 79 % of commits went straight to the main branch. An AI agent can forget a rule, so a written rule is not enough. Everything that matters now runs by itself when something happens (that is a gate), and every gate has a test proving it blocks what it should. A script checks that every written rule has a gate.
  • Every failure becomes a check (scar → gate): 160 real failures across my projects are catalogued, each with a rule id. I call them scars. Every checklist item cites the scar it came from (a rule taken from Google SRE). A postmortem ends with a new gate or a new scar, never with “be more careful”.
  • Proof, not claims: “done” means a command was run, it returned a result (exit code), and the state afterwards was checked. Every report ends with one word: VERIFIED, UNVERIFIED or FAILED. AI agents measurably overstate success (in one benchmark 75.8 % of reported successes had no evidence) and they give in when pushed back. The status word is the counterweight.

The life of one change

Nine checkpoints run by themselves at fixed moments: when I send a prompt, when the agent edits a file, and so on up to production. Nobody has to remember to check. Any red result stops the change on the spot.

LIVEpromptprotocol injectededitlint + typecheckcommandguard + no-verifysession endtier + tests+ reviewerscommit18 checks + secretspushno main + size+ cyclesCImutation + secretsmergecurrent merge-refonlyproduction200 from real domain✕ no silent skipping: quality gates are passed only with a log entry · risky steps only after a sub-agentreview · irreversible steps and the gates themselves only after my “allow ALLOW_…” in the chat · a routinepull request merges without my phrase only with proof of reviewNine checkpoints on the path of a change. I make the decisions; the checks make sure nothing is forgotten.
  • When I send a prompt: a hook (a script that runs when something happens) adds the working rules (the protocol) to the prompt. If the task is unclear, the agent asks questions instead of guessing. A fact needs an opened source first. Anything blocking the work goes in the first line of the report. And “how is it going?” never gets the answer “fine”: it gets the goal, the budget, the risks and the decisions that belong to the sponsor.
  • After every file edit, including edits made in the terminal: the changed file is checked for style and type errors (lint and typecheck). Errors go straight back to the agent that made the edit, in the same session.
  • At every terminal command: a guard blocks the dangerous ones: skipping the checks (--no-verify), rewriting or throwing away history (force-push, hard reset), recursive deletes outside build folders, merging pull requests from a script (the one exception is the merge script that demands proof of review), secrets typed into a command, and running a downloaded script straight in the shell (curl | sh). Since September it parses a command the way the shell would, so it also catches commands hidden inside sh -c, $(…), a heredoc piped into a shell, encoded PowerShell or tricks like rm${IFS}-rf. It guards the guards too: an agent can't quietly edit the hooks, the settings or a linter config, and a hook stops the AI reviewers from writing anywhere outside their own findings file.
  • When a session ends: the stop gate decides how risky the change is (the risk tier) from what actually changed: file paths and size, never from what the prompt says. It runs lint, types and tests on everything that changed. A session at tier T2 or higher cannot close until the required AI reviewers have checked it. The test verdict comes from the exit code: an import error or zero tests run is red, not “skipped”. Committing before the session ends does not lower the tier, and a review the aggregation script marks as incomplete does not close the session.
  • On every commit: 18 hard stops: conventional message; secret scan; base freshness (a clone on an unrelated history is blocked); duplicate literals in new code; a new dependency must exist on npm or PyPI and not be one typo away from a popular package; commented-out code; a new TODO without a ledger row; a personal-data column without a privacy inventory row; SQL migration lint (row-level security with both USING and WITH CHECK, definer hygiene, tenant foreign keys); GitHub’s own workflow parser on workflow files; a removed security step in CI.
  • At every push: nothing goes straight to the main branch. A change of more than 400 source lines is split into smaller ones. A new circular import (import cycle), or an import that breaks the layers declared in the architecture document, is blocked. If someone else committed to the same branch in the last 24 hours, you get a warning.
  • On the build server (CI) and at merge: the same checks run again, plus a secret scanner (gitleaks) and mutation testing on the changed files: the tool plants small bugs on purpose, and a test that catches none of them is theatre. 33 repositories have strict branch protection, generated from the names of the CI jobs. A merge script accepts a pull request only if its checks ran on the current merge result: once, green checks on an outdated base broke main in production.

Risk level (tier): measured from the change, never declared

The amount of checking is proportional, and that is enforced too. A prototype isn't nagged for architecture records. A change to payments cannot close without a security review.

  • T0: docs, copy, styles, assets: script checks only, with no AI tokens used.
  • T1: a single, isolated component or helper: code review is recommended.
  • T2: shared logic, API routes, edge functions, dependencies, CI or build configuration, or more than 150 lines. Required: code and operations review, plus UX review when the interface changes and data review when the schema changes.
  • T3: login and permissions (auth, row-level security, multi-tenancy), payments, secrets, SQL migrations, scheduled jobs, admin, or more than 600 lines. Required: code, security, data and operations review, then a verifier. The security, data and verifier roles run on the strongest model.
  • For each repository: a minimum tier, a lifecycle stage (prototype · poc · mvp · production) and an optional QA address. Moving to production requires a written production-readiness review in the same commit. With a QA address set, the QA reviewer becomes mandatory for T3 changes that touch a user interface.

Review like a software house, not like a chat

Each AI reviewer covers one speciality and starts with a clean slate (fresh context). A script, not a model, combines their findings: an issue blocks the change only if enough reviewers report it independently or the verifier reproduces it, and an issue without proof is dropped (the k-of-n rule). The verifier's only job is to disprove.

ORCHESTRATORSCRIPTSlint · types · tests · SQLcodesecuritydata · opsux · product · qaAGGREGATEk-of-n · needsproofVERIFIERreproduce or refuteFIX + RE-GATE✕ no chat between agents · a finding without an executed command does not countReviewers get only code that passed the scripts. A script merges findings. The verifier hunts for bugs, not confirmations.
  • Nine specialist reviewers: code, security, data, operations, UX, product, QA, a verifier and a devil's advocate (catfish). Each works from a checklist of at most eight numbered rules. Every rule has a command to run, a severity policy, a fixed output format (JSON schema) and a worked example of a false alarm it must reject. Reviewers only read; the main session makes the fixes.
  • Why a clean slate: whoever wrote the code, a human or a model, doesn't see its own mistakes. Reviewers start empty, get only the change and the task, and never see each other's findings. Agents that discuss drift toward the majority, even when the minority was right.
  • Decision councils with a devil's advocate: architecture decisions go through fixed steps: facts → positions → a mandatory dissenter → aggregation → a written decision record. Injected dissent is the one intervention shown to cut failures caused by quiet agreement in groups of agents.
  • Calibrated monthly: once a month the same bug is packaged two ways: a bare change, and the same change with a persuasive description. Six fresh reviewers then check it in random order. If they disagree, length or description is biasing them, and a human adjusts the checklist. The auditor never adjusts the instrument it measures.

Every failure becomes a check: how the system learns

A failure isn't closed when it's fixed. It's closed when it can't happen again without a script noticing. Each catalogued failure is called a scar.

incident orrepeated correctionroot cause:five whysscar +rule idgate + testthat it blocksevery rule hasa gateweeklyauditThe loop feeds itself: the weekly audit finds checks that stopped working, and each one becomes a new scar.
  • An outdated base broke main: two pull requests each passed their checks, but each was checked against its own older copy of the main branch. Merged together, they produced a broken workflow file on the main branch. Now: merge only when the checks ran on the current merge result, strict branch protection on 33 repositories, and a scar that records the exact mechanism.
  • A check that never ran: the push gate read its input twice and, for six days, silently did nothing. Now every git hook has a test that runs the whole script with real input, and the weekly audit checks that every gate still blocks what it should.
  • 45 copies of a company identity: spread across 11 files, and every one of them passed lint, types, tests and review. Now a duplicate check runs on added lines: the same text three times in two files is blocked. An escape hatch exists, and every use of it is logged.
  • An invoice that silently changed the seller: a default value (“?? default”) quietly swapped the seller after a profile was deleted. Now a silent fallback (SILENT-FALLBACK) on any field with consequences is a blocker in the code review checklist: when no match is found, the code must stop with a named error.

Update, October 6, 2026: a row for every area of system design

Version 1.5.0. A system design course covers 37 areas. Among them are queues (waiting lines for tasks), caching (keeping a ready copy of data for speed), scaling, databases and API design. Others are rate limiting (capping how many requests one user can send), real-time updates and search. I asked a question about PG, my set of rules and agents for writing code with AI. Does its architect, the agent that plans a project before any code is written, know these areas and walk through each one in every design? The October 6 audit showed that it knew 72% of the named options by name but could choose between only 46% of them. Nothing made it go through all the areas either. A second audit showed that the reviewers, the agents that check each code change, were overloaded too. This release changes both, and I tested it the same day on a real client application.

  • What the audit found: it counted 437 named options across the 37 areas. 46% came with a rule for when to choose them, 26% had only a name and 28% were not there at all. By area, 14 were covered, 21 partly and 2 not at all. The two missing ones were distributed consensus (how several servers agree on one answer) and probabilistic data structures (compact structures that give approximate answers to save memory). The design process forced a decision in 15 areas, partly in 11 and not at all in 11.
  • A row for every area: PG now has a list of the 37 areas in 12 groups. The first six are state and topology, data, identifiers, files and CDN (servers that deliver files from close to the user), consistency and asynchronous work (things done later, in the background). The other six are API contracts (the agreed way programs talk to each other), identity and access, resilience, observability (seeing what the system is doing), releases and disaster recovery, and cost and region. Every design ends with a matrix, a table with one row per area. Each row is marked DECISION, NOT APPLICABLE or NOT NOW and carries evidence. A NOT NOW row also needs a measurable signal that says when to come back to it. A script called sd-matrix-lint checks that the table is complete. The end-of-session gate (an automatic check before a session can finish) can block when a repository requires it.
  • Twelve agents instead of one architect: instead of one architect handling everything, there is now a council: one agent per group, twelve in total. Groups whose decisions are hard to undo or involve security get a stronger model. The decision cards, short notes the agents choose from, now list options with a rule for when to pick them, their cost and how they fail. Examples include caching strategies, rate limiting algorithms and backpressure (slowing the sender when the receiver cannot keep up). Others are SLOs (the promised level of service) and RPO and RTO (how much data and how much time you can afford to lose after a failure). The list ends with real-time features, RAG (answering from your own documents) and deployments.
  • Each reviewer reads its own slice: the second audit looked at one review round where every reviewer got the whole diff (the list of all changes in the code), 9,799 lines. In practice each one read about 10 to 30% of it. A CSS bug (CSS is the code that controls how a page looks) went past three reviewers. After two branches (separate lines of work on the code) were merged, two @media blocks (style rules for a given screen size) were left unclosed. As a result, part of the styling did not work on desktop computers. Only the tester (the step that clicks through the app like a user would) caught it. Now a script cuts the diff by department (security, data and so on), so each reviewer gets the part that concerns it. On a 2,400-line diff, security got 33% and data 14%. A second script runs 21 mechanical search rules and hands the hits to the reviewer to judge.
  • New checks that cost no tokens: tokens are the pieces of text a model is paid for, and these three checks use none. After every edit, one check makes sure every opening bracket in the CSS has its closing one. After a merge, every line from both branches has to survive. And the tester now warns about test paths that never click anything, because a green result there proves nothing.
  • Tested on a real client app: on October 6 I ran it on a client application built with Next.js and Supabase (a popular web framework and database service). The 12 agents filled all 37 rows (31 decisions, 5 not applicable, 1 not now) and listed 66 gaps with evidence. One of them: a scheduled daily data download had never run, because the schedule file was not on the main branch. I checked it, and it will start once the open pull request (a proposed change waiting to be merged) is merged. At least one gap was a false alarm. The sd-matrix-lint script returned 0 (pass) on the full matrix and 1 (fail) after one row was removed. The same app also went through a review round with slices and got 11 remarks. The verifier (an agent that checks whether a finding holds up) confirmed 2 more serious ones. The first was untested keyboard focus logic (focus is where your typing goes). The second: after adding a task, the focus landed on an empty page. Both were fixed the same day.
  • Honest limits: the matrix makes sure every area gets considered, but it does not guarantee that the decision written in it is a good one. The gaps the council lists have to be reviewed by a person, since at least one was wrong. The counters in the README (the project's description file), apart from the file count of 365, were not recalculated for 1.5.0. The public 1.5.0 release is in coding-higher-mind, the public repository of PG, merged on October 6 after the checks passed.

Update, October 5, 2026 (evening): the same rules for a much smaller context bill

Version 1.4.0. Every message to an AI model carries background text the model reads first (the rules, the memory notes), and it is paid for in tokens, roughly pieces of words. I measured that background for PG: about 26 thousand tokens before any work starts and about 1,300 more on every larger request, most of it rules the model had already read. The goal was to send much less while the model keeps following the rules just as well. Four reviewers and a verifier checked the release (17 findings, all fixed and re-checked).

  • The full rules once per session: Prompt-Guard, the part that adds the rules to my messages, now sends them in full once per session and afterwards only a one-line reminder. When the conversation is compressed or a new session starts, it sends them in full again, and whenever it is unsure what was already sent, it sends everything. A follow-up request about code went from about 4.9 KB of added text to 0.9 KB.
  • Memory that is current and still complete: the memory notes loaded at the start of a session used to show checkpoints from a week ago, because they were read from the oldest end of the file. Now the newest ones come first, unfinished work from other sessions is listed by title, and every shortened file keeps a list of its remaining sections, so nothing disappears from view. The start-up text went from 35 KB to 25 KB, and a budget check turns red if it grows back.
  • A stricter guard against switching off security checks: one gate stops an agent from quietly disabling the automatic security checks that run on every change. It raised a false alarm on a fresh project template, and false alarms teach an agent to work around a gate. Fixing that opened four gaps, which the security reviewer found: an unusual way of writing a step in the settings file, a whole job marked “continue on error”, the word “true” written in quotes or as an expression, and a pretend check that only appears in a comment or a printed line. All four are closed and covered by 16 tests.
  • GitHub's own vulnerability alerts feed the monthly check: Dependabot, GitHub's service that warns when a library a project uses has a known security hole, is now on in 45 of my repositories, with notices only on GitHub and no e-mail. The monthly vulnerability check reads these alerts alongside its own scan. Products in use get a fix through a pull request; the other repositories get one summary line in the log.
  • Honest limits: the savings are measured in kilobytes of text. Whether the model follows the shorter reminders just as well I will only know after a week of real sessions, and a scheduled check on October 12 looks at exactly that.

Update, October 5, 2026: pull requests merge without my phrase, but only with proof of review

Typing “allow ALLOW_MERGE” for every routine pull request had become the new bottleneck. So I rebuilt merging the same way as the exceptions: what decides is how risky the change is and whether there is proof, not whether I am at the computer. It has been switched on since October 4, and I accepted the remaining risk knowingly. A system like this needs time to mature. This is a version in the middle of that, not a finished product, and the limits are below.

  • What the merge script checks: every check green (the newest run of each one), not a fork, aimed at the default branch, the full list of changed files known. Then by risk tier: documentation only (.md, .txt, .rst) merges with no review; a normal code change needs a code review of exactly this diff, and a change to shared logic needs code and operations reviews. The review has to name the diff by its checksum, the verdict is recomputed from the findings and from the reviewers' own transcripts instead of the summary, and it is valid for two hours. Anything touching logins, payments, migrations, dependencies, CI or the agent's own instructions, and anything at the highest risk tier (T3), still waits for my phrase.
  • A second proof the agent cannot reach: the review kept on the agent's own machine could in principle be forged by that agent, and the security review found several ways. So a pull request can also get a review in CI: a clean runner, read-only tools, Claude's configuration taken from the base branch instead of from the pull request. The merge script reads that run through the GitHub API and accepts it only if the workflow, the PR number and every attempt on that commit are green. It keeps 90 days of verdict history, so rewording a change and rolling the dice again does not work on the same branch, and a deleted run shows up as a gap. It is set up in three repositories so far (I checked the workflow and the secret in each). In the pilot the reviewer stopped several changes with red runs and tends to be over-zealous, so I am watching it for a week before I trust its verdict more.
  • The stop gate no longer mixes up sessions: when two agent sessions work in one repository, a commit from the other session used to raise the risk tier of mine, or was ignored. Now a commit counts for the session unless there is positive proof that it was made during another session's tool run; when in doubt it counts. A new architecture check at T2 and above asks for a short decision record (at least five added lines) when a change adds a dependency, a table or a piece of infrastructure.
  • Design knowledge built into the reviewers: the architecture-advisor skill and the reviewer checklists now carry a set of decision cards for common building blocks (queue, cache, replica, CDN and so on): when it pays for itself, my default (usually “you don't need it yet”), what it costs and how it fails. Design before code now asks for a rough capacity estimate and a failure mode for every external dependency. The cards are my own distillation of system-design practice, and they ship in the public repository with version 1.3.0.
  • Automatic session notes, repaired: at the end of each session a hook now writes a short note into my Obsidian vault. The old version had been silently broken since September 22, because it read a field that no longer existed. The new one replaces it and has its own test, and the guard audit checks that it is wired in.
  • Honest limits: the CI review is one model, so it is a probabilistic proof that a crafted diff can try to talk around, and a reviewer can simply be wrong. A new branch with a new pull request and a cosmetically changed diff gets a fresh draw of the review. An agent with push rights can edit the workflow in an abandoned pull request and read the secret. No text rule in a command guard closes every way of swapping a reviewer. That is why the check must not be marked as required in branch protection (GitHub counts a skipped job as passed) and why there is a kill switch: one file turns phrase-only merging back on. Everything is written down in the repository's known-limits file.

Update, 2 October 2026: the agent approves itself, but only with proof

The “allow ALLOW_…” phrase protected things well, but it stopped the work. The agent would finish a change, a gate would say “diff too large” or “new TODO without a backlog entry”, and everything waited until I typed two words in the chat. I rebuilt it into three levels. Whether the agent can carry on alone now depends on whether the step can be undone, not on whether I happen to be at the computer.

  • Level A, the agent decides: ten quality gates (change size, TODO without an entry, commented-out code, repeated literals, module boundaries, commit format and similar). The gate still shows the problem. The agent fixes it or deliberately moves on, and every such pass goes into the log with the gate's name.
  • Level B, approval after review: deleting files, a hard reset, cleaning the repository, a dependency missing from the registry and a new column with personal data. The agent asks sub-agents for a review (security and code, plus the data department for personal data), fixes what they find and only then grants itself an exception. The agent only gets it when the code matches what was reviewed. It lasts 15 minutes and covers exactly one command, only in that repository.
  • Level C, still only me: changes to the gates themselves, pushing to main, merging a pull request the merge script will not take on its own (see the update of 5 October above), secrets in a commit, weakening CI or quality settings, force-push, deleting remote repositories. These can't be undone or they switch off other checks. If the agent could approve changes to its own gates, levels A and B would stop meaning anything.
  • Proof, not a claim: the first review of the new mechanism showed that the proof of review could be faked with two empty files. Now what counts are the transcripts of real reviewer sub-agents. Every finding a reviewer recorded has to stay in the result, and one review never gives a second exception, even if the folder is copied. A command that tries to leave the repository (changing directory, a variable, a shell wrapper) doesn't get the exception.
  • Four review rounds before shipping: the code, security and operations departments went through the change, and the data department joined for the last two rounds. Every round found a new way around the rules: through a variable, through a shell wrapper, through a PowerShell quirk, through a one-line script. Each one is fixed and has its own test. The gate test set now has 104 cases, and mutation tests noticed each of the 25 rules being switched off.
  • The clock slip: while I was working, the system clock briefly jumped to 2022. The file holding my exception got a date four years in the past, and the cleanup deleted it as old, even though it was valid for another half hour. The cleanup now reads validity from the file's content, not from when it was written.
  • Kill switch and limits: one switch brings back the old mode, where every exception needs my phrase, and it is protected the same way as the gates themselves. The limit is the same as before: text rules stop an agent taking shortcuts, not someone deliberately forging transcripts or building names out of pieces. I've written it down as a known limit.

Update, 26 September 2026: PG next to 35 open-source projects

I checked what everyone else is building. AI agents went through 35 public repositories with a similar idea (hooks, gates, AI reviewers). They compared each one with PG layer by layer, and every point cites a file and a line number. The result was a list of 23 things other people did better. I shipped the important ones the same day.

  • The design was reviewed before any code: before the first line was written, an independent security reviewer took the plan apart. It found 15 problems, 3 of which would have broken the review process itself. All of them were fixed in the design.
  • Tested on real history: every new rule in the command guard ran against 40,824 real commands before it was allowed to block anything. They come from 2,853 agent transcript files, main sessions and sub-agents. The false alarms this caught were fixed before rollout. One example: text being written into a file was read as a command. Checking one command takes at most 16 ms.
  • The checks are tested too: a set of 102 cases (84 for the command guard) has to pass in full, and every rule needs one case it blocks and one it lets through. On top of that, mutation tests switch off each of the 24 rules in turn and check that the tests notice. They noticed 24 out of 24.
  • Reviewed after it went live: four AI review departments (code, security, data, operations) and a verifier went through the finished change. They found 26 problems; 17 are fixed and confirmed. The verifier keeps finding more elaborate variants of the same few bypasses, such as a drive letter mapped onto the settings folder. Text rules can't win that race against someone who is hunting for holes. That needs isolation at the operating-system level, and I've written it down as a known limit.
  • Context survives compaction: when a long session gets compacted, a hook first saves a snapshot: my last instructions, the files edited and the last answer. After compaction it goes straight back into the context, so the agent doesn't have to rebuild the state from a summary.
  • Exceptions come only from me (until 2 October, see the update above): an agent used to be able to add an ALLOW_…=1 switch to a command and walk past a check. Now an exception exists only after I type “allow ALLOW_…” in the chat. It lasts 30 minutes and three uses, and every use goes into the log. The same phrase echoed in an agent’s reply or in a sub-agent’s report unlocks nothing.
  • Protecting the protection: the hooks, git gates, settings and gate tools are sealed with checksums. A change outside a deliberate window shows up at the start of every session and in the weekly audit. The rule that denies access to SSH keys, AWS credentials and the secrets vault also holds in the mode that never asks for permission. I checked that in a live session.
  • A monthly look at the competition: a routine compares those 35 repositories with the saved snapshot. The script runs without AI; a model only steps in when something changed in the code that does the checking, or a new, fast-growing project shows up.
  • Not done yet: a nightly check (a canary) that runs the agent on a pinned tool version and confirms the gates still block, checking CI status when a session ends, and testing skills under pressure. All of it is planned, and I'll add it here once it runs.

Two AI tools, one system

  • Claude Code: the terminal tool for work on code: this site, the production apps behind MAS Group, Flyt and the garage system, and their CI. It connects to GitHub, mail, the browser and the desktop through MCP servers (connectors). This is where the hooks and git gates apply.
  • Claude Desktop (Cowork): the tool for operations: mail triage, documents, research, outreach, design, and the routines that keep the system itself running. It has no hooks, so the working rules travel as text inside every prompt.
  • 105 skills: versioned instruction packages that tasks are routed to: 50 in the coding tool, 55 on the desktop. A router matches each task by fixed rules (deterministically) and announces which steps it will run before work starts.
  • Choosing the model: the strongest model is kept for hard reasoning and for the security, data and verification reviews at T3. Routine tool work and lookups go to cheaper models. Computing capacity is a budget, and the system spends it on purpose.

VERIFIED · UNVERIFIED · FAILED: the status word that stops the AI telling you what you want to hear

Every substantive report ends with one of three words. It's the smallest part of the system and the easiest to take elsewhere, especially to Claude Desktop (Cowork), which has no hooks, so the prompt is the only protection.

  • VERIFIED: the claim comes with its proof, quoted rather than described: the command and its exit code, the HTTP status, the line from the test output, the diff, the path to a screenshot.
  • UNVERIFIED: the work is done but the proof is missing. The report says exactly what is missing, how to check it, and what would change the conclusion. A report with no status word counts as UNVERIFIED.
  • FAILED / BLOCKED: what happened, word for word, without softening. Anything blocking the work (missing access, a decision, data or a secret, or a broken tool) goes in the first line of the report, never at the end.
  • Why it exists: when a user pushes back with a wrong claim, models agree in about 58 % of cases. They predict 61–77 % success and achieve 22–35 %. The longer a conversation runs, the more the model mirrors your framing and your confidence. The status word forces every claim to show evidence or admit it has none. And “are you sure?” makes the model re-check the evidence instead of politely changing its answer.
  • In Claude Desktop and scheduled tasks: every background prompt gets a seven-line verification block: don't assume the prompt is true; you may refuse and report failure; restate claims as neutral questions; no success without evidence; attack your own result before reporting; re-check when challenged; end with the status word. The block is in the repository, ready to paste.

Passwords and keys the AI never sees

VAULT99 secrets · self-hostedBRIDGEpasses to the processTARGET PROCESSdeploy · API call · CIchat ✕code ✕logs ✕The agent can use a password or key it can never read.
  • A vault on my own hardware: 99 API keys, tokens and logins are stored in a vault on my own hardware. A bridge passes them straight into the process that needs them, as environment variables. The values never appear in the chat, the code or the logs. The command guard blocks a secret typed into a command, and the scans at commit time and in CI keep it that way.

Routines: the part that runs while nobody is typing

Five routines in the coding runtime, 35 defined on the desktop, 10 of them enabled. Deterministic script first, model only on findings; anything that could “find itself work” is off by design.

PR reviewerweekday mornings
Watchdogevery 2 hours
Config backup + restore scriptdaily
Sessions → memory notesdaily / Mon + Thu
Safeguard statusweekly
Weekly system reportSunday
CVE watchmonthly
Reviewer calibrationmonthly
Repo cleanermanual only
  • PR reviewer (weekdays): reviews open pull requests across all my repositories, the way a senior engineer would. Mechanical fixes land as separate commits with proof attached. Design and security findings stay as comments for me to decide.
  • Safeguard status (weekly): 33 script checks that the quality system itself is still connected: hooks are registered, gates still block what they should, which escape hatches were used and why, and which routine started but never finished.
  • Vulnerability watch (CVE, monthly): first a script scans the dependencies of the live products: no AI tokens used when nothing is found. Fixes are limited to patch and minor updates, and always arrive as pull requests with evidence, never as direct pushes.
  • Watchdog (every two hours, Claude Desktop): finds routines that are overdue or crashed mid-run, retries them and fixes what it can. It sends a phone notification only when it can't. If a usage limit interrupted work, it resumes from a saved checkpoint.
  • Backup with a restore script (daily, Claude Desktop): a full copy of the agent configuration plus a generated script that restores it on a new machine in one click. A backup that was never restored is not a backup, so the restore script is part of the backup.

Memory, measurement and the retrospective that rebuilt the system

  • Persistent memory: a local base of 1,900+ notes is loaded at the start of every session. Each session writes its own notes when it ends, so the next one starts from a saved checkpoint instead of from scratch.
  • Telemetry, not feelings: every skipped check is logged with a reason. Scripts mine 2,853 agent transcript files and the git history of 39 repositories: corrections per session, repeated tool errors, fixes made within 24 hours of the previous commit to the same file, and the files that change most often (churn hotspots). The numbers decide what becomes a gate.
  • The retrospective that mattered: twelve repair loops in one session were traced to their causes. Three had the same defect: the system claimed a check it didn't physically have. The answer was structural: every gate got a test proving it blocks, and checking that every rule has a gate became a script.
  • The same method, applied to me: my code-reading practice is built like the rest of the system, daily, verified, public: github.com/kamiljan11/code-reading-quest.

Does it hold at scale?

  • The honest answer: one function is easy. A 200,000-line system, with dependencies between files and unwritten assumptions about its architecture, is where consistency drifts. So the assumptions are written where a script can read them: every repository's architecture document has a machine-readable block of layers and forbidden imports, and the push gate blocks a new import cycle or an import that breaks the layers. Old cycles only warn. Historical debt is never cleaned up automatically.
  • Checking the impact before an edit: a changed file that more than 40 other files import is flagged. The design step before any code asks: who calls this, what else reads this data, and which risk tier the change falls into.
  • Readable by a person, not only by a machine: the tool index and the system map are generated from each tool's own header; a tool without a description shows up as debt. Every style decision is judged by one question: can a senior engineer who has never seen the repository run it in 15 minutes, find the place to change in 15 minutes, and understand why, without reading my transcripts?

Install it yourself

The whole system is a public repository, cleaned of private data and portable. You only need Node 20+, git and Claude Code.

  • git clone https://github.com/kamiljan11/coding-higher-mind.git
  • node install.mjs --dry-run # shows the plan, touches nothing
  • node install.mjs --yes # copies into ~/.claude, merges hooks, appends the CLAUDE.md block, wires git hooks
  • What the installer promises: it never overwrites a file you changed (the other version lands next to yours). It merges your settings instead of replacing them. It adds its rules between markers, so an update replaces only that block. Git hooks are optional. It ends with a self-test: the green output is the proof, not the installer's word.
  • What you get in the public version today: 10 hooks, 3 git gates, 39 tools with 11 test suites, 9 reviewer departments, the doctrine with 160 scars, a repo template with CI and parsed boundary blocks, 7 coding routines and 7 desktop routines, an uninstaller, and docs with the diagrams from this page.
  • One honest note: the working rules have a full English version (PG_LANG=en, which the installer sets from your system language), and every block message includes an English BLOCKED line with its escape hatch. The doctrine and the reviewer checklists are still in Polish (the model reads them fine), and the README is in both languages.

Open the repository →

Honest limits

Foundation models via API: I do not train or fine-tune them. Reliability is proven at SME scale (dozens of repositories, one owner), not hyperscale. Reviewer departments cost tokens (roughly four times one review for T2 and eight to ten for T3), which is why zero-token gates run first. Merging without my phrase rests on one model's review (in CI or local), which is a probabilistic proof; see the update of October 5 and the known-limits file. Some gates depend on the repository having what they check, and skip with a logged reason when it does not. The point of this page is not that the system is finished: it is that the failure modes of working with AI are engineered against, in the open, instead of being wished away.

This page went through the process it describes: an agent drafted it, it passed the gates above, and I reviewed and published it. The numbers came from commands run on the day it shipped, not from memory.