We’ve overloaded CI
I can’t remember working professionally without a continuous integration (CI) server. Over time, there are always efforts underway to make CI faster and, at the same time, to cram more tests and gates (linters, validators, architecture checks, etc.) into its run.
I’m starting to think we should question the second part. Not the value of CI itself, but rather the habit of turning every check we can think of into a gate.
I can’t remember working professionally without a continuous integration (CI) server. Over time, there are always efforts underway to make CI faster and, at the same time, to cram more tests and gates (linters, validators, architecture checks, etc.) into its run.
I’m starting to think we should question the second part. Not the value of CI itself, but rather the habit of turning every check we can think of into a gate.
What CI was for
Continuous integration was a response to a very specific pain: Several teams build a large piece of software. Each team works to spec, builds its features, writes its tests, and moves on. Separately. Then comes integration and the pieces don’t fit.
The fix was to integrate the parts early and often, so that every team always knew whether the parts, put together, actually worked. Combine that with agile and lean, and you get software that grows from a small MVP you can hopefully validate toward something with more and more functionality—and that keeps integrating, i.e. keeps running properly, the whole way. Continuous deployment (CD) takes the last step: as soon as something is ready, it goes out to the customer.
CI/CD is a practice I rely on every day.
The build server became a thing-doer
The job to be done is integration. The vehicle is something else: an arbitrary thing-doer. Modern CI systems are containerized. Each step is a stateless, reproducible container, and the output of a command in that container tells us whether a test, or a suite of tests, passed.
Once you have an arbitrary thing-doer, you can, of course, make it do other things.
Some of those expansions are natural. The original bar was that the integrated system compiles. Then, that its tests pass—that it behaves correctly. Then, perhaps, that it meets performance requirements or some other -ility. All of these still answer the original question: does the whole thing work? And more: does it work well?
And then it expanded more: What other things can one enforce as a gate in CI? Linters, data checkers, validators, architecture checkers, and many, many more.
What makes teams do this? If a structural rule gets broken once, they all know they’ll never get around to fixing it. So they might as well encode it as a permanent check and gate every merge to main on it. A gate settles the question of whether you have to do it.
Linting doesn’t break anything
Here is the curious part. If CI only existed to make sure the integrated system works, a linter wouldn’t belong in it.
Unlinted code does not cause incidents because it is unlinted. It just looks a certain way that someone decided we don’t want. Several versions of the same code are all fine: they compile, they pass the tests.
I looked at a repository I work in. It has more than 30 different kinds of checks. I sorted them by one question: would the system stop working if this check failed? Half the checks would not break the app on failure.
Then I asked what share of the load on the build system those checks cost. Almost none. These checks tend to be static. Static checks are fast by construction, and the slow ones tend to get fixed: Packwerk took minutes on large code bases until Alex Evanczuk rebuilt it in Rust as pks, and now it takes seconds.
So the cost isn’t compute. The cost is that they are gates. And with more and more small fixups handled with minimal human interaction, I believe we have to ask whether there’s a better way to make sure our systems are not just correct, but also conform to the non-breaking constraints we’ve set for them.
The browser that didn’t need a green build
About eight months ago, a team had hundreds of agents implement a web browser. At the start, they effectively told every agent to land clean, green commits. Given our history of CI, that was perfectly natural. Of course you produce a green build!
Then they realized the swarm as a whole went faster when they dropped that requirement:
When we required 100% correctness before every single commit, it caused major serialization and slowdowns of effective throughput. Even a single small error, like an API change or typo, would cause the whole system to grind to a halt. Workers would go outside their scope and start fixing irrelevant things. Many agents would pile on and trample each other trying to fix the same issue.
This behavior wasn’t helpful or necessary. Allowing some slack means agents can trust that other issues will get fixed by fellow agents soon, which is true since the system has effective ownership and delegation over the whole codebase.
Step back and it makes sense. The goal was a browser that works, measured against browser specifications, which are remarkably complete. They wanted to prove the agents could build the whole thing. They were not trying to prove the agents could build it while keeping a test suite green the entire time.
Didn’t they just end up with a broken browser? Not necessarily. Their suggestion: one more agent that regularly takes snapshots and does a quick fixup pass. Ninety-nine agents stop being responsible for the build — one agent becomes responsible for nothing else.
Where the gate stops paying
How much does gating have to cost before 99 + 1 beats 100?
Say every gated agent loses some fraction of its speed to keeping the build green. A hundred gated agents produce the work of 100 × (1 - slowdown) agents. Ninety-nine ungated agents plus one fixer produce the work of 99. With a hundred agents, the gate only pays for itself if it costs each agent less than 1%.
None of this neatly leads to “delete all your lint checks.” And, of course, if your product is live in front of customers, you should need to keep the checks that validate it actually does what you claim it does.
But, I believe, this points at an opportunity: CI doesn’t have to be, dogmatically, the place where every check runs. Rather than a question of principle a la “this must run in CI,” it is an optimization problem. And you can run this analysis for every gate that doesn’t actually enforce that the app works.
You can’t optimize what you can’t run
Here’s the catch. You can only exploit this optimization if you can actually run cleanup tasks like this.
So here is my challenge. If you don’t run any cleanup agents automatically—nothing that responds to commits, nothing that runs on a schedule—you have no way to even find out where your trade-off lies. There is no mechanism to explore how much you gate changes on their way to production versus how much you fix out of band: style, taste, architecture, all the things that don’t break anything.
If your application has any complexity, build that capability. Then start finding out which checks should leave the build, stop being gates, and become things an automated system fixes.
Chatter
These are webmentions via the IndieWeb and webmention.io. Mention this post from your site: