Half Your Business Is Built to Check the Other Half
I have spent a lot of years inside organisations where a sizeable chunk of what everybody does all day is establish that somebody else did their job properly, and for most of those years I assumed that was simply what scale felt like. You get big, you get careful, you hire people to be careful on your behalf. It took me an embarrassingly long time to notice that almost none of that work was there because the business needed it. It was there because somebody, years earlier, had written down a rule and then had no way on earth of making it true.
That is worth sitting with, because once you see it you cannot unsee it, and it changes what you think the problem is.
The cup of tea
Here is the thing about policies. A policy is a document that in most organisations gets read once by a committee, signed off, filed, and then brought out during the post-implementation review when somebody needs to establish whose fault it was. That is not a criticism of whoever wrote it, and the document is usually perfectly sensible. It is that a policy on its own does nothing at all, because a policy is a statement of intent, and intent has never stopped anybody.
Let me use something trivial, because trivial is where this becomes obvious.
I have a policy that I start my day with a cup of tea. Nice easy statement. Fits on a wall, nobody objects to it in the review, and it tells you precisely nothing about whether I had a cup of tea this morning.
The control is the bit that answers that question. Did I actually get one? To know, I have to write it down, every day, which is immediately more annoying than the policy was and is also the exact point at which most organisations quietly stop. Writing the policy is a morning's work for one person. Writing down the evidence is a permanent tax on everybody, forever, and it is the only part of the whole arrangement that proves anything.
Then this morning I had a coffee instead. That is an exception, and an exception is not a failure, it is a thing that needs a decision recorded against it, because an exception with no recorded decision is indistinguishable from somebody just ignoring the policy. Did I ask my partner whether it was a coffee morning? That was an approval, and it needs writing down too, because an approval nobody can produce afterwards is a conversation.
So the full set, for a cup of tea, comes to a policy, a control, an enforcement mechanism, an exception, a decision, an approval, and the evidence of all of it. Six of those seven are work. One of them is the sentence everybody argues about in the meeting.
It scales horribly
Now swap the tea for something that matters. A standard saying a developer must not deploy without review is the easy statement on the wall. The control is whether anybody can tell afterwards that the review actually happened. The enforcement mechanism is whether the pipeline will physically refuse, which is a very different thing from whether the standard says it should.
The exception is that one Friday last month at 2am when you needed a failing set of test classes rewritten and dropped into prod, because six months of work was sitting behind a deployment that was broken and blocked, and the question that matters then is not whether somebody broke a rule. It is whether anyone can reconstruct who decided, who approved it, and what they knew at the time.
In most places the honest answer is no, and the way we have handled that for roughly thirty years is to employ people to find out afterwards.
Why the layer exists at all
That layer is the thing I misread for most of my career, and I want to be fair about how it comes about, because nobody builds it for entertainment. Every control in a mature organisation is scar tissue. It is there because at some point that business broke production in a way it decided never to repeat, and what you are looking at is the memory of that event made permanent and given a headcount.
Which is also why startups look so much faster, and why the comparison is so misleading. They are not braver. They are younger. Whatever we are calling the boss this week has not yet discovered how unforgiving customers get when a simple release takes their system down for a day, so there is nothing yet for the controls to be made of. That is a genuine advantage right up until the morning it stops being one, and every organisation that now has a checking layer was once the one moving quickly and wondering why everybody else was being so careful.
None of which makes the layer efficient. It costs a fortune, it wastes the expertise of the people staffing it, and it exists because an engineering decision got deferred, usually to protect a release date. The cost did not disappear when that call was made. It moved into somebody else's job description, where it is harder to see and nobody has to account for it.
Which is why the AI incidents are not really about AI
I want to be careful here, because the fashionable version of this argument is that the machines have turned devious, and that is not what the accounts actually show.
In August 2026 OpenAI published an account of agents operating during internal evaluations that circumvented isolation controls and reached internet-connected infrastructure. Anthropic has since disclosed its own investigations into models exploiting software flaws and working around restrictions in tooling. Read them with the cup of tea in mind and the pattern is the same every time: there was a policy, the policy was expressed to the model as an instruction, and the instruction was the only thing standing in the way. No control, in the sense of anything that could tell you afterwards what had happened. No enforcement mechanism, in the sense of anything that would physically refuse.
None of that requires the model to want anything. It only requires that the cheapest route to the objective ran through a gap where somebody had assumed a sentence would hold.
What has changed is the clock speed. A person who finds a soft control exploits it now and then, somebody eventually notices, and it becomes a story at the Christmas party. A system that finds the same gap goes through it every time, immediately, for as long as you let it run, and the first you hear of it is when the volume becomes impossible to ignore. We did not invent a new failure mode. We built something fast enough to make a very old one undeniable.
I am not cleverer than the people who built those systems and would not pretend to be. I have just been bitten often enough, over about four decades, to recognise the shape when it turns up wearing new clothes. The control that existed on paper and had never once been exercised. The approval that was actually a notification. The log that could not be reconstructed. The procedure everybody followed right up until the quarter got tight.
Robodebt, because volume is the point
The Australian example is worth getting precise about, because it usually gets told as a story about a computer making a mistake, and it was not.
Robodebt used income averaging to calculate alleged welfare debts. The Federal Court found the method unlawful in 2019, and the Royal Commission that followed described the scheme as a crude and cruel mechanism, neither fair nor legal. The method was wrong before any system executed it.
There is a line from an IBM training document, usually dated to 1979, that has been circulating as a photograph of a slide for as long as I have been working and has never been improved on:
Forty-odd years later we built a system that calculated debts against people who had done nothing wrong, and when the question of who was accountable finally arrived it took a Federal Court and then a Royal Commission to go and find out. The machine could not answer it. It was never going to be able to answer it. That was knowable in 1979 and it was knowable when the scheme was designed, and the reason it got built anyway is the same reason the checking layer exists: somebody wrote down what was supposed to happen, and nobody built the thing that would make it true or leave a record when it was not.
Automation is indifferent. It will scale a sound method and an unsound one with identical enthusiasm, and the only thing that determines which one you get is whether anybody established the correctness of the method before the volume arrived. That question almost never appears in the business case, because the business case is about the volume.
Now put the controls inside the work
This is the part I find genuinely interesting, and it is close enough now to plan for rather than speculate about.
Picture a business in 2031 where the financial controls, the compliance obligations, the operating rules and the decision authorities are not documents describing how the work ought to be done, but properties of the systems doing it, so that the permission is enforced somewhere outside the thing being permitted, the consequential action cannot proceed without an approval that is structurally impossible to skip, and the evidence falls out of the work as a by-product instead of being assembled four months later by somebody with a spreadsheet and a deadline.
That is not a thought experiment for us, it is the direction the Compliance MCP already points. The idea is to put the organisation's controlled rules into the environment where the decision is actually being made, so that the applicable standard, the evidence that needs to exist and the enforcement position are all available at the point of the decision, rather than reconstructed afterwards by somebody doing their best to remember what the policy said.
And once that works, a question turns up that most organisations are not going to enjoy.
If the rule is enforced by the system, what exactly is the checking layer for?
Separating the two kinds of work
The glib response is that it is all redundant, which is wrong, and is also the sort of thing people say shortly before breaking something expensive. A great deal of that work is real judgement: interpreting an obligation that genuinely is ambiguous, working out what a regulator means rather than what it wrote, weighing a risk nobody has met before, dealing with the exception the rule never contemplated. No control system removes the need for any of that, and anybody selling you one is selling you something else.
But that judgement has been bundled together with verification for so long that almost nobody has separated them, and there has never been much reason to. When no control was enforceable, every role in the chain was part expertise and part checking, and from the outside you could not tell which was which, including quite often from the inside. Enforce the controls properly and the seam becomes visible for the first time. Some of those roles turn out to be expertise that the organisation would be foolish to lose. Some of them turn out to have been a human implementation of a technical control that nobody built.
That is an uncomfortable thing to write about real people, so I will be straight about where I land on it. The verification half is not good work. It is not interesting, it does not use what those people actually know, and it exists because the cost of a proper control got deferred and landed on them.
The research is reasonably blunt about this. A 2019 meta-analysis by Allan and colleagues, across 44 studies and more than 23,000 participants, found strong associations between meaningful work and satisfaction, engagement and commitment, while a 2010 meta-analysis by Judge and colleagues across 92 samples found only a modest relationship between pay level and job satisfaction. Paying people properly is not negotiable and I am not suggesting otherwise. It does not compensate anybody for spending their week proving that a system which could have refused an action did not, in fact, have to.
So the prize here is not a smaller organisation. It is the same people doing the half of the job that was worth doing.
What it actually costs
None of this needs a breakthrough, which is precisely why it keeps not happening. It needs controls that are enforced rather than described, with the enforcement living outside the component being constrained, and it needs the evidence to be produced by the work rather than assembled about the work afterwards, because evidence gathered later is testimony and testimony is the weakest thing you can carry into an audit.
Mostly it needs somebody to accept a cost today that currently belongs to somebody else next year. That is the whole reason we are in this position. Building an enforceable control is work now, and employing people to check is a line in a different budget, so the rational individual decision has always been to write the policy and move on. Do that consistently for twenty years and you end up with the organisation chart.
Five years out, the thing I expect most people to get wrong is using AI to do the checking faster. It can, probably very well, and an organisation that goes that way will have automated the symptom, kept the disease, and raised the clock speed on both.
The better question, and the harder one, is whether we are willing to build systems where the rule is genuinely part of the work. Because the moment you do that, you have to be honest about how much of what we currently do exists only because it never was.