ENGARPAI Talk to sales
Governance practice

What can your AI system evidence today?

Compliance scores cannot be audited. The question that can be is narrower, harder and far more useful: what is on the record for this system, right now, and where does the record run out?

EAENGARP AI Engineering9 min read

The short answer

  • An AI system is governed only as far as its weakest link, so the useful unit of measurement is a chain of records, not a score.
  • Nine links, from registered to release gate. Each one either has a record behind it or it does not.
  • Nearly every portfolio breaks in the same place: between controls scoped and controls effective.

The wrong question, asked confidently

Ask most AI governance tooling whether you are compliant with a framework and it will answer with a number. Seventy-three per cent. Amber. Three of five stars. The number is arithmetic performed on judgements that are not commensurable: a missing model card and an unreviewed adverse decision do not belong on the same axis, and the weights that put them there are a vendor's opinion rather than a regulator's.

The deeper problem is that a score cannot be audited. No auditor accepts a percentage in place of a record. They ask what you did, when, who decided it, and what you can show them. A governance programme built to move a number will optimise for the number, and then discover during its first real audit that the number was never the deliverable.

A system is only as governed as its weakest link, and a score is designed to hide exactly that.

The nine links

Replace the score with a chain. Each link is a record that either exists or does not, and the sequence matters because the later links depend on the earlier ones being true.

  1. 01RegisteredThe system exists in one inventory, with an owner, a business unit and a citable code. Shadow systems are the single most common reason a chain cannot be drawn at all.
  2. 02Risk assessedA completed assessment producing a tier from a stated taxonomy, with the deciding rule visible. A tier nobody can reconstruct is a number, not an assessment.
  3. 03Controls scopedApplicability computed from risk tier, autonomy and data classes, so the control set follows from the system rather than from whoever filled the form in.
  4. 04Controls effectiveNot implemented, not attested: effective, on evidence. This is the link that separates a governance programme from a documentation exercise.
  5. 05Evidence attachedVersioned, hashed, reviewed, expiring. Evidence that can be silently overwritten is worse than none, because it looks like proof.
  6. 06Independently testedTested by someone who does not own the outcome. Self-testing is how a control passes for two years and fails the week an auditor arrives.
  7. 07Findings resolvedA failed test becomes a finding with an owner, a due date and a retest. Closing it requires a new run, so the failure history survives.
  8. 08ApprovedA named human accepting a named risk, with separation of duties enforced. A system owner approving their own risk acceptance is not an approval.
  9. 09Release gateA deterministic verdict that publishes its reasoning check by check, and that later change can invalidate but never edit.

Where the chain breaks first

In practice the break is almost always between link three and link four: controls scoped, but not evidenced as effective. The reason is structural rather than cultural. Scoping is a one-off modelling exercise that a small team can complete in weeks. Effectiveness requires evidence from control owners and independent testing on a recurring cadence, which is a standing obligation on people who do not report to the governance function.

This is why the portfolio view matters more than any single system's page. Thirty-one of forty-eight systems clearing link three and only fourteen clearing link four is not a data quality problem to be cleaned up before the board sees it. It is the finding.

A test worth running

Pick your most business-critical AI system. Without asking its owner, try to produce: the current risk assessment, the list of controls scoped to it, the last independent test result, and the approval that let it into production. Time yourself. That duration is your real governance posture, and no score will improve it.

Five honest states, and no sixth

If a score is out, something has to take its place at the top of a report. Five states are enough, and each is defensible in a room with an auditor in it.

The only five things a governance record should ever claim about a requirement.
StateWhat it asserts
CoveredMapped to controls that are effective, with current evidence and an in-date independent test.
Partially coveredMapped, with some controls effective and others not. The gap is enumerated rather than averaged away.
Evidence pendingControls are implemented and the requirement is mapped, but the artefact that proves it is missing or expired.
Not assessedHonestly unknown. The most valuable state on the list, and the one every scoring model destroys.
Not applicableOut of scope for this system, with the reason on the record and a reviewer attached to it.

Notice what is absent. There is no state that means compliant. Compliance is a legal conclusion drawn by a qualified person against a specific regulation and a specific set of facts, and a platform that asserts it is taking a judgement it is not entitled to make, on behalf of someone who will be asked to defend it.

What to do on Monday

None of this requires a platform to begin. It requires deciding that the chain, not the score, is what you report.

  1. Write the nine links down and agree them with internal audit before you instrument anything. If audit will not accept the chain, the tooling cannot save it.
  2. Register every AI system, including the ones built in a business unit without telling anyone. An inventory that excludes the uncomfortable entries is not an inventory.
  3. Pick one internal control library and map frameworks onto it. Adopting a second standard should add a view, never a second backlog.
  4. Make effectiveness depend on independent testing rather than self-attestation, and accept that your first honest report will look worse than your last dishonest one.

Questions we get asked

What does it mean for an AI system to be governed?

That it can evidence each link of the chain: registered, risk assessed, controls scoped, controls effective, evidence attached, independently tested, findings resolved, approved, and past a release gate. It is governed only as far as its weakest link, because one broken link fails the gate regardless of the links after it.

Why is a compliance score the wrong output?

Because it averages judgements that are not commensurable, using weights that are a vendor opinion, and because no auditor accepts a percentage in place of a record. Mapped, covered, partially covered, evidence pending and not assessed are all defensible; seventy-three per cent is not.

Where does the assurance chain usually break first?

Between controls scoped and controls effective. Scoping is a one-time modelling exercise; effectiveness is a standing obligation on control owners and independent testers, so most portfolios have far more controls scoped than evidenced as effective.

Can a platform declare an organization compliant?

No. Compliance is a legal conclusion drawn by a qualified person against a specific regulation and set of facts. A platform can state what is mapped, covered and evidenced. Asserting compliance transfers a judgement it is not entitled to make to a tool that will not be in the room.

Does AI governance require an internet connection?

It should not. Inventory, risk, controls, evidence, workflow, approvals and audit can all run inside client infrastructure with no egress. Licence verification, framework content delivery, model providers and telemetry are the four dependencies that need offline equivalents, and each has one.

All field notes

Published 2 September 2026 by ENGARP AI Engineering

Keep reading

Run the test on your own portfolio.

We will walk the nine links against one of your real AI systems, in ninety minutes, with an engineer rather than a demo reel.

Talk to sales