The complete guide to the five-level scale: what each level means, what separates one from the next, and the objective gates that decide when a team has actually cleared one.
Every company we talk to has bought AI. Almost none of them can tell you what it changed. The licences are approved, the tools are in every IDE, and the release calendar looks the way it looked last year. When you ask why, the answers are anecdotes: this developer is faster, that team likes it, someone automated a script.
Anecdotes are not a baseline, and without a baseline you cannot decide anything. So we built a scale. The AI Delivery Levels describe how a team actually delivers with AI, not which tools it owns, and they do it in a way that produces a number a board can act on.
There is only one dimension: who starts the work. Not the number of licences, not the size of the budget, not how enthusiastic the team is.
Read that list again and notice what is missing: tools. A team can run the best model on the market and sit at DL2 forever, because the model has no idea what it is working on.
Roughly two thirds of teams are at DL2, and the shape of it is always the same. Everyone has a licence. Everyone uses it their own way. Nobody has written down what the AI needs to know about the project, so every answer starts from zero: generic code, the wrong library, naming that does not match, an error-handling pattern nobody uses here.
The individual gain is real and small, somewhere around 1.25 times. It is also invisible at team level, because it does not compound. When the person who figured out the good workflow goes on holiday, the workflow goes with them.
DL3 is the first level where the gains belong to the team instead of to individuals. The mechanism is not a better tool, it is a shared, versioned context that the AI reads before every task: your conventions, your architecture, your constraints, maintained by the people who know the system.
Once that exists, the AI stops guessing. It matches your patterns on the first try, flags conflicts between services before review, and suggests the tests somebody would otherwise have forgotten. And critically, a correction made once is a correction the whole team inherits.
This is also the hardest jump in the framework, and the classification rules make that explicit.
We score six core phases separately: requirements, architecture, implementation, testing, review and documentation. CI/CD, release and monitoring matter for delivery but do not decide the level. Then two different kinds of threshold apply.
That asymmetry is deliberate, and it produces the result teams argue with most: five phases at DL3 and testing at DL2 is a DL2 team. One weak phase holds everyone back, because an agent inherits the weakest link in the chain, not the average.
The other rule people underestimate: a phase can only get one level ahead of what it depends on. Implementation cannot outrun the specifications and the architecture it is built from. You cannot have AI generating complete features from specs when no specs exist.
This is why buying a coding tool so rarely moves the number. The tool lands in the one phase that was already the strongest, and the phases upstream of it stay exactly as they were.
The multipliers are ranges rather than promises, and they depend entirely on execution quality. DL2 sits around 1.25 times, and it stays individual. DL3 is where team-wide gains start to compound. DL4 is where delivery capacity changes shape, because the unit of work moves from the prompt to the request. DL5 is real but early: the tooling is months old, the governance frameworks are being invented in real time, and most of what exists is not production-ready at scale.
We would rather be honest about that last part than sell it. Most of our clients should be aiming at DL3, with DL4 as the horizon.
Score your six phases honestly, find the lowest one, and ask why it is lowest. In our experience the answer is almost never the tooling. It is that nothing written down is precise enough for an AI to work from, which makes specification quality the most common bottleneck in the whole framework.
A five-level scale describing how a software team delivers with AI: DL1 manual, DL2 individual and reactive AI use, DL3 shared context where the AI understands the project, DL4 agentic delivery across the whole lifecycle, DL5 autonomous operation at organisation scale. The single dimension separating them is who starts the work.
Because DL3 and above are completeness thresholds: every one of the six core phases has to be at that level. DL2 only requires two of six. One weak phase holds the whole team back, which is intentional, because agents inherit the weakest link rather than the average.
Not usefully. An agent inherits whatever context you hand it, and at DL2 that context lives in scattered heads and one-off prompts. Without the shared context of DL3, a fleet of agents simply reproduces your undocumented assumptions faster.
The six core delivery phases: requirements, architecture, implementation, testing, review and documentation. CI/CD, release and monitoring support delivery but do not decide the classification.
A structured assessment. We score the six phases with your team, place you on DL1 to DL5, name the one bottleneck holding the rest back, and put it in writing. It takes 30 minutes and costs nothing.
Want to know where your teams actually sit on the scale, and which single phase is holding the rest back?
Book the free assessment →