DevSecOps · Engineering Standards · AI

The AI Can Clean Up 60,000 Feature Flags. That Doesn't Make 60,000 Feature Flags Good Engineering.

  • DevSecOps · Engineering Standards · AI
  • Q3 2026
  • blog
A migration path with two branches, which is what a feature toggle actually is

There is a particular kind of technology story that makes me incandescent, and it is not because the technology is bad. Usually it is genuinely clever, built by people who are very good at their jobs. What sets me off is the framing: we take a problem created by years of accumulated engineering practice, build an impressive system to clean up the consequences, and then tell the story as though the cleanup itself were evidence of engineering excellence.

A recent InfoQ story about DoorDash describes a multi-agent LLM system built to clean up stale feature flags across its estate, and the numbers are genuinely arresting. DoorDash reportedly manages more than 60,000 feature flags across roughly 623 repositories, creates about 2,300 new ones every month, and has identified more than 1,000 that are stale. The agentic system investigates a flag, works out how it is wired through the codebase, makes the changes, runs validation and produces a pull request for engineering review.

I want to be fair about this, because the work is not trivial. The agents are doing something genuinely difficult to automate: the dependency-injection patterns involved mean removing a flag is rarely a simple Boolean deletion, and DoorDash reports that a single flag can touch five to twenty files including tests, with relationships between the flag and the application logic that are sometimes semantic rather than syntactic. That is real engineering and I would not pretend otherwise. But there is another way to read the same story, which is to ask why a construct that is supposed to be temporary became a population of 60,000, and why we are quite so impressed that a machine can now clear up after us.

A feature flag is not configuration

At AppGenie the rule is deliberately simple. A runtime switch is one of exactly two things: a feature toggle, which is a migration path, or an operational control, which is an executable path responding to an external factor. The distinction matters because it determines whether the thing is allowed to survive.

A feature toggle exists because the system is temporarily between two behaviours, so it has a lifetime by definition. Its eventual state is not "false", it is not "inactive", and it is certainly not "we have not touched it in a while". Its only permitted end state is removal. An operational control is a different animal: it exists because the system genuinely needs to respond at runtime to something outside itself, such as dependency health, entitlement, licensing or regulatory state, and it is allowed to be permanent precisely because that external factor is permanent.

The test for telling them apart takes about ten seconds. Remove the switch, commit the system to one branch, and see what breaks. If the system can no longer respond to an external factor, it was an operational control. If the only consequence is that the system now commits to the new behaviour, it was a feature toggle, and you have just done what should have happened months ago.

We do not leave that distinction to interpretation after the fact, which is the whole point. It is written down in the AppGenie Engineering Coding Standard, ENG-CTRL-01, and section 5.13 is deliberately uncompromising: a feature toggle cannot be environment-injected, cannot hold different values in different environments, must default to the new behaviour, and must be introduced with a removal date and a linked work item. Removing the toggle while leaving the superseded code path behind is non-compliant, because that is not a migration, it is a mess with better branding. And past its declared removal date, the build fails.

This is not a recommendation, and it is not a retrospective cleanup process. It is a control.

The declaration itself is mundane, which is rather the idea. Both kinds of switch have to say what they are at the point they are introduced:

feature-toggle: <Name> removal:<YYYY-MM-DD> ref:<WORK-ITEM>
operational-control: <Name> factor:<external factor> owner:<role> ref:<REGISTER-ENTRY>

The interesting part of the DoorDash story is not the AI

The less comfortable question in that story is not how the agents work, but how any organisation accumulates a problem at that scale in the first place. I want to be careful here, because this is not an argument that DoorDash developers are careless. I have no basis for saying that, and I would not pretend I could walk into a 623-repository estate and understand it better than the people who built it. I could not, and that is actually the point.

I am not smarter than every developer at DoorDash, or at any other large engineering organisation. I have simply been bitten by a lot of things over roughly four decades, which is long enough to watch the same failure modes come round again wearing different clothes. Configuration that quietly became permanent. Temporary migration code that became architecture. Environment variables that changed application behaviour without anyone noticing. Database switches nobody could explain. Dependencies nobody knew were there. Code paths that were definitely dead, right up until somebody discovered production depended on them. Documentation that described what the system was supposed to do rather than what it actually did. Controls that existed on paper and had never once been exercised.

And then, eventually, somebody asks the question that should have been asked at the beginning, which is what exactly are we running. That accumulated experience is what goes into the AppGenie controls. Not genius, and not cleverness. Experience, most of it acquired the expensive way.

Do not believe me. I learnt this lesson from Doug.

If you think I am being unnecessarily paranoid about controls, do not take my word for it. Go and read Doug Seven's Knightmare: A DevOps Cautionary Tale. His account of the Knight Capital incident stays with you because the individual decisions involved are far less interesting than the system they eventually produced. The lesson is not that developers are stupid; it is almost the opposite. People make reasonable decisions in the context they can see, and systems retain those decisions long after the context that justified them has disappeared. That is exactly why controls matter.

And if you want to know what AppGenie actually requires of feature toggles, do not take my word for that either. Go and look up ENG-CTRL-01 section 5.13 in the Compliance MCP controlled standards, because the lesson should not live in this blog post, it should not live in my head, and it certainly should not depend on whether the developer implementing the next migration happens to have been bitten by the same problem before. It belongs in the control system.

This is why we built the Compliance MCP

The AppGenie Compliance MCP exists because putting a standards document in a folder and asking an LLM to read it is not enough. An AI assistant can retrieve a document, summarise it accurately, and produce an extremely convincing answer, and none of that means the resulting work is controlled. What it means is that the model read something.

The MCP puts the organisation's rules and controls into the environment where the AI is making decisions. The point is not to make the model smarter, because it already has more general knowledge than any of us. The point is to give it the specific accumulated experience of this organisation, expressed as controlled, applicable rules. That is a fundamentally different proposition from asking a model:

"What does ISO 9001 say about this?"

The question we actually want answered is the one with our context in it:

"Given our standards, our controls, our environment and this particular engineering decision, what is permitted, what is required, and what evidence must exist?"

This is not just a feature-flag rule

Section 5.13 does not sit in isolation, and that matters more than the specific rule about toggles. The same standard defines explicit non-compliance conditions: a runtime switch introduced without being declared as either a feature toggle or an operational control, a toggle past its removal date, a toggle removed while its superseded code remains, and an operational control that is unregistered, unowned, or has never been exercised.

It also defines what counts as evidence, which is where most good intentions quietly die. Source-control history, linked work, pull-request review, test results, static analysis, security scanning and traceability all count. Developer statements do not, and that sentence exists in the standard for a reason that anyone who has sat through an audit will recognise immediately.

The standard then names the enforcement points across development, local validation, code review, static analysis, testing, CI/CD, release readiness and audit review, and failure blocks progression unless an approved exception exists. That is the bit that turns a piece of good engineering advice into an actual control, and it is the difference between a standard and a suggestion.

This is the same problem as the SBOM

This connects directly to something else we have written about, in Can You List What Is In Your Software?, which asks the deceptively simple question of what is actually in a piece of software. An SBOM answers part of that, and it is the part the industry has decided to care about, largely because regulators got there first.

But software carries a second inventory that we tend to ignore entirely, which is behaviour. What behaviours can this system execute? Which decisions can change at runtime, which are permanent, which are temporary, which are environment-dependent and which are controlled from outside? Why does each one exist, who owns it, when does it expire, and what evidence proves it has ever been exercised?

You can know every library in your application and still have no idea which of several materially different behaviours production will execute tonight.

A dependency inventory without a behavioural inventory is half an answer. And the gap between the two is not a documentation problem that a wiki page will fix. It is a control problem.

The dangerous phrase is "we can clean it up later"

This is one of the oldest traps in the trade, and it is always phrased reasonably. We will remove that later. We will document it later. We will rationalise the configuration later, consolidate the environments later, update the dependency later, and put the evidence together before the audit.

Sometimes later genuinely arrives. Usually later is the point at which the person who made the decision is no longer thinking about it: the system has moved on, the people have moved on, the repository has changed, the operational environment has changed, and the temporary thing has quietly become part of the architecture. Nobody decided that. It simply happened while everyone was busy.

Which is why AppGenie controls tend to be deliberately uncomfortable. They are not designed to make the developer feel good about the commit. They are designed to make the failure mode difficult to reach.

Compliance should encode experience, not paperwork

This is also why I do not think compliance should be a collection of PDFs that developers consult when somebody from governance asks for evidence. That model puts compliance downstream of engineering, arriving after the decisions have been made in order to find out what was decided. We are interested in the opposite arrangement, which is to put the control where the decision is made.

If a developer introduces a migration mechanism, the control should apply there. If an AI assistant is writing the implementation, it should apply there. If an agent is preparing a release, it should apply there. And if a system claims that something is an operational control rather than a feature toggle, there should be a rule requiring that distinction to be made and evidenced rather than asserted. The result is not "AI compliance", which is a phrase I would happily never hear again. It is something much more practical: engineering experience that a machine can apply.

We are not claiming to be smarter

This is probably the most important thing I can say about what we are building. I do not think I am smarter than the developers at DoorDash, and I do not think the AppGenie team is smarter than the engineering teams at Google, Microsoft, Amazon or anyone else operating software at enormous scale. What we have is considerably less glamorous than that.

Four decades of watching reasonable decisions accumulate into unreasonable systems. Four decades of seeing the same problem resurface long after the original context evaporated. Four decades of learning that "temporary" is one of the most dangerous words in software, and that the best control is almost always the one applied at the point where the decision is made rather than the one applied after the consequences have piled up. And four decades of seeing what happens when nobody had that control. Doug's Knightmare is one of those lessons. There are plenty of others, and most of them are less famous and a good deal more ordinary.

You do not need to be smarter than the people who made the original mistake. You need to have learned from it, and then you need to put that learning somewhere it cannot be forgotten. Experience becomes standards, standards become controls, controls become machine-readable, and machine-readable controls become available to the AI through the Compliance MCP. At which point the model gets something considerably more valuable than another paragraph of prompt: it gets the benefit of the scars.

The goal is not to make AI clean up our mess

None of this is an argument against what DoorDash has built. AI will absolutely make it cheaper to clean up technical debt, that is valuable, and their work demonstrates it clearly. But cleaning up debt is not the same thing as not creating it, and we are at real risk of confusing the two precisely because the cleanup is so much more impressive to watch. There is a difference between:

"Our AI can remove stale feature flags."

and:

"Our engineering system makes an expired feature flag a build failure."

The first is an impressive capability. The second is a control, and the distinction gets more important as AI writes more of the software itself. Once a machine can produce code at enormous speed, we cannot keep relying on somebody eventually finding everything it created, because there is no longer any realistic "eventually" in which that happens.

So the accumulated experience of engineering has to go into the environment where those decisions are being made. That is what the Compliance MCP is for. Not to make the AI cleverer, not to replace developers, and not to pretend that having a compliance catalogue makes software safe. It is there to make decades of lessons available at the exact moment a decision is being made, because the most valuable engineering experience is usually not knowing how to build something. It is knowing what you will regret building if nobody stops you now.