DevSecOps · Azure DevOps · Pipelines

Who Watches the Watcher? The DevSecOps Dirty Little Secret

  • DevSecOps · Azure DevOps · Engineering Governance
  • Q3 2026
  • blog
Who watches the watcher - the DevSecOps dirty little secret

The bit nobody in this industry volunteers

Here is the dirty little secret of DevSecOps consulting. A great many of the firms who will happily audit your pipelines, write you a maturity assessment and tell you your release process is a liability have build pipelines of their own that would not survive the same review. Not because they are dishonest. Because their pipelines are internal, nobody is paying for them, and the standard you enforce on a client is a service line while the standard you enforce on yourself is an overhead.

So this post is our own pipelines, with the actual numbers, including the part that is still not finished. If you are going to write something called "who watches the watcher" you do not really get to leave that bit out.

The thousand-line YAML problem

Anyone who has worked in Azure DevOps for a while has met the pipeline that has become a small application in its own right. A thousand lines of YAML, forty inline Bash and PowerShell blocks, and the accumulated sediment of every deadline anybody has ever hit. It usually starts as one reasonable file and grows because inlining a script is always faster than putting it somewhere sensible.

The problem is not the length. It is what length hides. A hard-coded account number in an inline script at line 780 is invisible in review, because nobody reviews line 780 of a YAML file. They review the fifteen lines of the diff. Enormous pipelines are where hard-coded ARNs, embedded subscription IDs, silently skipped test steps and the odd credential go to live quietly and undisturbed.

A long pipeline is not complex because the deployment is complex. It is long because nobody has ever had to justify a line of it in isolation.

What we actually run

Our approach is dull and it is the whole point. The consuming pipeline should be a manifest, not a program. It declares what is being deployed and which template does the work. Anything with actual logic lives in a shared, versioned template that is reviewed once and used everywhere.

Here is the shape of it, counted this morning rather than remembered:

Measure Actual
Deployment pipelines in the services repo21
Total YAML across all of them1,738 lines
Average pipeline length82 lines
Longest single pipeline178 lines
Pipelines that call a shared template18 of 21
Shared templates in the core repo29
Total template code behind them4,448 lines

Read those two totals together, because the ratio is the argument. There is more logic in the templates than in all twenty-one pipelines combined, which is exactly right. The complexity has not disappeared, it has moved somewhere it can be reviewed properly, tested once and changed in one place. Twenty-one deployments, and the average one is eighty-two lines of "here is what I am and here is the template that knows how to ship me".

The practical benefit shows up when a control changes. When we needed every Lambda deployment to record a deployed-artefact entry, that was one template, once. Not twenty-one pull requests and a spreadsheet tracking which teams had got round to it.

Standards nobody reads are not controls

Templates fix structure. They do not fix behaviour, and behaviour is where the interesting failures are. For that you need gates: small, boring, fast scripts that fail the build.

We run six of those as static checks across the estate. They ban environment variables as a configuration or secret store, hard-coded credentials, a specific HTTP client that broke a monitor at the edge, unpinned package feeds in two ecosystems, and untracked feature toggles. On top of those sit another eight domain gates on the trading codebase, each asserting a structural property that a design decision requires, such as one module never being permitted to reference another.

The environment variable one has the most instructive history, so it is worth being specific. A deploy wiped a Lambda's environment variables and took a production service down. The relevant engineering standard already banned using environment variables that way. It had banned it for a while. The ban was in an annex, and the annex had not been read.

So the standard was rewritten as a script that fails the build. Its own header says it plainly: it is the binding control behind the policy, so that the ban does not depend on anyone reading a standard.

The detail I would steal

That gate ships with a self-test, and the self-test replays the incident. You can run it in isolation, in a second, with no credentials and no pipeline:

$ python validate-no-env-vars.py --selftest

[PASS] EdgeLock.cs      expected_violation=True  got=True   (code-read)
[PASS] Program.cs       expected_violation=True  got=True   (code-read)
[PASS] activate.sh      expected_violation=True  got=True   (infra-write)
[PASS] defaults.json    expected_violation=False got=False  (none)
[PASS] defaults.json    expected_violation=True  got=True   (infra-write)
[PASS] ok-suppressed.cs expected_violation=False got=False  (suppressed)
[PASS] bad-suppressed.cs expected_violation=True got=True   (suppression missing ref)
[PASS] clean.cs         expected_violation=False got=False  (none)

SELFTEST: OK

Two things in there matter more than the pass count. The third case is the exact change that caused the outage, so the gate can demonstrate on demand that it would have caught the thing it was written for. That turns "we added a control" into something you can actually verify, which is the difference between a control and a claim.

The seventh case is the one I would steal outright. Every gate needs an escape hatch, because occasionally a rule genuinely should not apply. Ours allows an inline suppression, but the suppression must carry a reference to the decision or gap that justifies it. A suppression without a reference is itself a violation, and the self-test proves it. The escape hatch is policed by the same gate it escapes, which is the only version of an escape hatch that survives contact with a deadline.

So who does watch the watcher?

We do, and it is not flattering, which is rather the point.

That environment variable gate was built the day the incident was resolved. It passed its self-test immediately. What took considerably longer was wiring it into the build templates so that it actually ran on every deployment. Our corrective action register recorded the gate as built with its CI wiring still outstanding, and a later monitoring incident recorded the same thing more bluntly: the custom code gates existed and no pipeline invoked them. Written, proven, and not yet plumbed in.

Most of that has since been closed out and the majority of those gates now run in the pipelines. But the gap between "we built a control" and "the control runs on every build" was real, it was measured in weeks rather than hours, and it existed in the estate of a firm that does this for a living.

The only reason I can tell you the specifics is that it was written down at the time, in a register, with an owner and an open status, rather than quietly fixed and forgotten. That is the actual control. Not the gate. The habit of recording the gap between what you have decided and what you have deployed, and leaving it visibly open until it closes.

One of our own scripts carries the lesson as a comment, which is where these things usually end up:

A gate that is not in the pipeline is documentation.

What to take from this

  • Make the consuming pipeline a manifest. If it contains logic, that logic belongs in a template. Eighty lines you can read beats a thousand you skim.
  • Put the complexity where it gets reviewed. More lines in shared templates than in all your pipelines combined is a healthy ratio, not a smell.
  • Turn standards into scripts. Anything that depends on somebody having read an annex is not a control, it is a future-scheduled failure.
  • Give every gate a self-test that replays the incident. It is the cheapest evidence you will ever produce and it stops the gate rotting silently.
  • Police the escape hatch with the gate. Suppressions must justify themselves or they become the new default.
  • Audit your own estate on the same terms as a client's. You will find something. Write it down with an owner instead of fixing it quietly.

None of this is clever. Short pipelines, shared templates, small scripts that fail the build, and an honest register of the distance between your standards and your reality. The last one is the one nobody wants, and it is the only one that tells you whether the rest of it is true.

On the numbers. Every figure above was counted from our own Azure DevOps pipeline and template files at the time of writing (Q3 2026), not estimated. Counts move as the estate changes. The incidents referenced are our own and are described here at the level of pattern; the specific incident, decision and corrective-action records behind them are internal.