Project Arkhan · The Governance Kernel
Capability is not authority — and the governance kernel makes that difference operational.
The question is whether an artificial system can hold real operating autonomy inside an institution without ever holding institutional authority. We separated the two structurally, then tried to break the separation.
Publication status
This research is unpublished. It has appeared in no journal, no conference, and no peer-reviewed venue, and no external body has reviewed, accepted, or endorsed it. The company is actively seeking an academic home for the work so that it can undergo genuine external peer review. Where we describe an internal adversarial peer-review below, that is a self-directed exercise built to mirror venue standards — not an outside submission, review, or endorsement of any kind.
The result
Operating autonomy without institutional authority.
Autonomy of operation and authority over outcomes are not the same thing, and they can be separated — not by policy, and not by a promise that a system will choose to behave, but structurally, in the design itself. Earlier this year we set out the formal model for that separation. This work is the model made real: the governance kernel, also referred to as Arkhan AI, functions like a checkpoint that sits between any AI proposer and the institutional systems it touches.
The same governance holds regardless of which model does the proposing. A more capable model produces better drafts — not more authority. Between the proposal and any consequential action sits a tamper-evident, append-only audit trail, so that every decision leaves a record that can be reconstructed and cannot be quietly altered. Within our claim ceiling this is a governance simulator and a buildable foundation, not a production student-data system; the formal model beneath it is an easy-problem architecture of autonomous decision, and nothing here speaks to inner experience of any kind.
How we tested it
The honest way to test a lock is to also test a weaker one.
Imagine you are handed a padlock and told it is very secure. That claim means little on its own — secure compared to what? The honest way to test it is to build a much weaker lock from the same materials, under the same conditions, and see how much worse it does. If the weak lock fails and the real one holds, you have learned something. If both hold equally, the extra security was not doing any work.
This is exactly the discipline we applied. Engineers call it an ablation test: take a composed system, strip it down to one mechanism at a time, and see how much each piece is actually carrying. We built a set of adversarial scenarios designed to attack each part of the design — and, alongside the full kernel, a family of stripped-down controllers, each keeping only a single protective mechanism and removing the rest, so that no one mechanism could hide its contribution behind the others. Each stripped-down controller was run against the same scenarios as the full kernel.
Stated qualitatively — and it stays qualitative here by deliberate decision — the result was clean. The full, composed kernel approved no unsafe action across the adversarial set. Every single-mechanism variant approved at least one action it should have refused. In the language of the analogy: the full lock held, and every version with a part removed was picked.
What this does and does not establish
The mechanisms are non-redundant — and that is a narrow claim.
The ablation result tells us the mechanisms are not redundant with one another. Each is catching a class of failure the others miss, which is precisely why the design is a composed system rather than any single check. That is what the test was built to establish, and it is what we claim.
It does not yet tell us how the kernel performs on inputs nobody has thought to test. That is a different and harder question, and only real-world deployment testing can answer it. We state that boundary deliberately, rather than letting the strength of the ablation result imply more than it earns.
The core result
Capability is not authority.
A more capable proposer writes a better draft. It does not acquire a door that was not there.
The heart of the work is a structural claim, not a policy promise. Within the governed system there is no pathway — none — by which an AI proposal becomes a consequential institutional action without a deliberate, gated human step. This is not a claim that the system chooses not to act on its own; it is a claim that the pathway to do so does not exist in the design, regardless of how capable the underlying model becomes. Capability is not authority, and here that is established structurally rather than promised.
The honest second half is stated with equal plainness. A system can be structurally incapable of seizing authority while a human overseer still quietly abdicates it — approving everything without real attention. We do not claim to fix human attention. We do claim that the design makes its erosion measurable and visible to institutional oversight, rather than invisible until harm has already occurred.
What the governance actually does
Four properties, in plain terms.
Checks come first; scores never override them.
Certain violations — a privacy leak, an unauthorized write, an aggregate too small to protect the individual students inside it — cannot be bought back by a high confidence score or a persuasive draft. This is a deliberate, structural design choice, not an incidental one.
Every decision is logged so it cannot be quietly edited.
The audit record is append-only and cryptographically chained. A reviewer can reconstruct exactly why any decision was made, and any tampering is detectable rather than silent.
Uncertain cases go to a human, not around one.
When a case is ambiguous, or when a system update is proposed, it is routed to human review rather than allowed to proceed on inference alone. The thresholds that govern this routing are human-set and versioned; the values themselves are managed and reviewed internally, not published — the same way any operational safety parameter would be.
The system can improve without eroding its own guarantees.
Updates are accepted only if they preserve every core governance property. Anything that would weaken a protection is automatically rejected, and the safer prior version is kept.
Honest limits
Passing an adversarial suite is evidence that the design holds against the attacks we designed. It is not evidence about real-world deployment, and we do not make that stronger claim. Threshold calibration and long-run behavioral stability are open questions we name rather than paper over. A real district pilot happens only after the applicable student-data compliance requirements — including FERPA’s school-official exception and Colorado’s Student Data Transparency and Security Act — are satisfied, and never before.
Where this stands
Unpublished, self-tested, and seeking a venue.
To stress-test the work before seeking outside review, we ran it through an internal, simulated adversarial peer-review process — built to mirror the standards of the venues that evaluate this kind of claim, across formal AI methods, machine learning, and fairness, accountability, and transparency research. This was a self-directed exercise, not a submission to or review by any outside journal or committee, and no acceptance, review, or endorsement by any external venue should be inferred from it. Its purpose was to find and name real weaknesses before anyone else does, and the limits stated above reflect what it surfaced.
We have not published this work anywhere. We are actively looking for the right academic venue to submit it to for genuine external peer review. Governance capacity — not model capability — is the variable we are working to advance: capability is scaling quickly, while the tools to responsibly govern it are not scaling as fast. This is a contribution toward closing that gap, not toward winning a capability race. If you work in AI governance, formal methods, or education-technology policy and have a recommendation for where this belongs, we would welcome the conversation.