In short: The AI Engineering Production System is a model for redesigning engineering around AI agents: six stages from Direction to Observe & Improve, set inside an organisational frame, staffed by agents and people, and running on one layer of context and control. In McKinsey's 2026 survey, 30% of leaders said AI had lowered team productivity.
Why do most AI redesigns stall?
In May 2026 McKinsey surveyed 334 product and engineering leaders about AI in software development. Only a quarter of the directors and above reported meaningful acceleration, which McKinsey defines as more than a quarter of their teams at least doubling productivity. Thirty per cent said team productivity had fallen.
What separated the groups was redesign. Organisations that embedded AI and also changed roles and ways of working became top accelerators 40% of the time. Those that embedded AI without changing roles got there 27% of the time, and those still experimenting 17%. These are self-reported results from a consultancy survey, so treat them as a direction of travel.
Redesign is now the consensus. McKinsey's own list has four themes: rewire workflows, redesign roles, build verification and AI operations that keep pace, and invest in change. Each one is right. A list of themes still leaves a CTO with three questions: what exactly to redesign, in what order, and how to know it worked.
The previous piece in this series answered the third, with validated business value per unit of human intervention. Constraint Migration answers the second, because the order follows the constraint. This piece answers the first by laying out the AI Engineering Production System, the model the rest of this work hangs on.
What is the AI Engineering Production System?
The AI Engineering Production System is a model of how a technology organisation turns intent into validated change once AI agents join the workforce. It has four parts: a frame that sets focus and structure, a spine of six stages that every change passes through, a workforce of agents and people spread across those stages, and a layer of context and control beneath them. It is drawn to show where work waits, because that is where the constraint sits.

Figure 1. The AI Engineering Production System. The frame changes yearly, Direction monthly or quarterly, the flow with every change; the layer under it is built once and used by every stage.
The frame: Shape. Enterprise focus, organisation design, funding and decision rights. No change passes through it, so it sits around the spine as the structure the flow runs inside. It still holds constraints of its own, and the model has to be able to show them.
The spine: six stages. Direction chooses what to work on. Intent turns that choice into something an agent can act on and a test can check. Build produces the change, Verify & Release decides whether it ships, Production lands it, and Observe & Improve proves it in use and sends the evidence back. Every stage has a capacity, a queue, an owner and a measure.
The workforce: agents do, people decide. Agents take on more of the doing at each stage, densest in Build today. People keep the judgment at every stage.
The layer: Context & Control. Context is what agents know. Control is what they may do. Checks are how their output is proven at each stage. The layer is built once, owned as a platform and used everywhere.
The spine is the delivery lifecycle in the Methods dimension of MAMOS, plan, design, implement, deliver and improve, rewritten for a workforce that now includes agents.
How is this different from an agentic SDLC?
The past year has produced a crowded shelf of lifecycle models: the agentic SDLC, EPAM's Agentic Development Lifecycle, McKinsey's agentic product development life cycle and several spec-driven variants. Most take the familiar phases and ask where an agent can help in each. EPAM's adds a useful first step that maps responsibilities between people and agents.
Those models are task maps. The AI Engineering Production System is a constraint map.
Once each stage carries its capacity, its queue, its owner and its measure, the question changes. You stop asking where an agent could help. You ask where work waits, who owns that wait, and what it will move to once it is relieved.

Figure 2. Same six stages, two maps. Only the constraint map shows what to redesign next.
Lifecycle models skip that last question. Relieve one stage and the constraint reappears at the next least scalable one, which is Constraint Migration. A model that cannot show where the constraint sits cannot tell you what to redesign next.
The constraint map also has one objective. Every design choice in the system is judged by whether it raises validated business value per unit of human intervention. Without a target, a redesign tends to optimise whichever stage complains loudest.
What changes at each stage when agents join?
Direction and Intent. Upstream is where most AI programmes look last. The previous piece described the patterns that hold the constraint before a line of code exists: too many initiatives at once, and initiatives that cannot ship without each other. Agents make both worse, because they make starting work cheaper.
Intent is the hinge of the system. McKinsey's bar for a good requirement is that an agent can act on it, an engineer can challenge it, a test suite can validate it and a reviewer can judge it, all without a follow-up conversation. Add the outcome the change should move, declared before build. That declaration is what Observe & Improve checks later.
Build and Verify & Release. Build is where agents are densest and where most AI budgets go. Its constraint is rarely the speed of writing. When code becomes cheap, architecture becomes expensive, because agents meet every ambiguous boundary at volume and tend to resolve it by adding code.
Verify & Release is the first stop the constraint reaches, and the easiest place to see it. Faros AI's 2026 telemetry showed median review time up about 441%. McKinsey's survey shows the same imbalance from another side: respondents reported 11.8% average time savings and only 6.2% less rework. The checks themselves belong in the layer, so every stage proves its own output. The release decision stays a stage, because that is where work waits and someone signs.
Production and Observe & Improve. At La Redoute we brought commit-to-production under ten minutes and cut new-service creation from days to minutes, as the CNCF case study records. That pipeline was sized for human output. Agent volume is a different load, and pipelines and environments need resizing before agents arrive at scale.
Observe & Improve closes the loop in two places. Health evidence returns to Intent, so the next specification improves. Outcome evidence returns to Direction, so the portfolio learns which bets paid. Many organisations wire the first loop and leave the second loose, which is how delivery metrics improve while the roadmap stays the same.
Where do agents and people sit?
Agents enter the workforce stage by stage, close behind the constraint. Each round relieves one stop and adds a population that needs context and control. McKinsey reports that leading organisations already give agents scoped mandates, permissions and escalation paths. The unit of design becomes a team of people and agents with defined responsibilities.
People hold the judgment at every stage: what the business meant, whether a change is safe, and who answers for it. Agents add to that load rather than relieve it. That is why Constraint Migration ends at human judgment, and why the signature metric divides by human hours.
Ownership should follow where work waits. At La Redoute, once the platform was in place, engineering took on shift-left architecture, business metrics and level-two support, while operations moved to owning a self-service platform. Each shift brought ownership closer to the queue it affected. Agents make the same move more urgent, since queues now shift faster than job descriptions.
Why does the frame belong in the model?
The model runs on three clocks. The flow moves with every change. Direction has usually moved quarterly; as agents shorten delivery, a monthly cycle becomes realistic, and the right cadence is the one that keeps pace with the constraint. The frame of budgets, org charts and decision rights usually moves once a year. Constraint Migration can now relocate the constraint within a quarter, which makes the frame the part most likely to be behind.
Some constraints live only in the frame: teams cut across the flow of value, funding attached to projects rather than products, decision rights that send every cross-team change into a meeting. No stage-level fix reaches them. At La Redoute, planning had coupled teams to each other. The answer was architectural and organisational together: APIs and events let teams ship independently, and order-to-shipment time fell from days to two hours.
Redesigning the enterprise is a larger subject, and part of the Systemic CTO work. In this model the frame has a narrower job: it names the structure the flow runs inside, so a constraint sitting there can be seen.
Why are context and control one layer?
Every agent at every stage needs the same three things: context about the work, limits on what it may do, and checks on what it produced. Built team by team, they become dozens of private arrangements nobody can audit. Built once, they become a platform every stage draws on.
McKinsey's survey points the same way. Organisations whose AI operations function held five or more mandates, such as governance standards and agent performance monitoring, reported more impact on quality and time savings. McKinsey also estimates that spend on tokens, compute and AI operations can reach 20% of existing labour costs, which makes cost one of the controls.
The layer is where the second front of the Absorption Gap sits. The Second Absorption Gap showed agents outrunning the context organisations can govern. Practitioners now call part of this work harness engineering, and Context Engineering Is the Wrong Altitude argued that the organisational job is larger than any of its practitioner names. In this model, the harness is how engineering builds part of the layer.
How do you run it?
A model becomes a management practice through four habits.
Put one ratio at the top. Validated business value per unit of human intervention is the objective. Stage metrics are the diagnostics that explain why it moved.
Name an owner for the constraint. Someone should be able to say where it sits this quarter and where it is heading.
Design one stop ahead. Before relieving a stage, check that the next one can absorb what it will receive.
Review each part on its own clock. The flow is managed continuously, Direction at least quarterly and monthly where delivery allows, and the frame at least twice a year while migration outpaces it. The layer is run as a product with its own roadmap.
For the redesign of each stage, MAMOS works as the checklist: what changes in methods, architecture, management, organisation and skills. A redesign that touches one dimension usually finds the constraint waiting in another.
Where should a CTO start?
Pick one high-value flow. Map it across the six stages and measure the queue time in front of each. The longest, fastest-growing queue is the current constraint. Then look behind it for the frame: a team boundary, a funding line or a decision right that keeps it there.
Redesign that flow end to end, from Direction to Observe & Improve, and set a baseline for the ratio with the 30-day test in The Absorption Gap. One redesigned flow teaches more than six separate stage programmes, because it shows where the constraint goes next.
The map before the redesign
Every AI programme will redesign something this year. The AI Engineering Production System exists so that it redesigns the right thing first, and can tell afterwards whether it worked.
If you redrew your engineering organisation around where work waits, how many of today's team boundaries would survive?
Read further
The thesis in full: Redesigning engineering production for the agentic era
The model's home page: the AI Engineering Production System
What the system is optimised for: What Is Your AI Engineering System Optimising For?
The mechanism behind the order of redesign: AI Won't Replace Your Engineering Bottleneck. It Will Move It.
The input-side front: The Second Absorption Gap
The design lens: MAMOS: Architecting the Software Production System for Quality at Speed
FAQ
What is the AI Engineering Production System?
It is a model for redesigning a technology organisation once AI agents join the workforce. A frame of enterprise focus and organisation design surrounds a spine of six stages: Direction, Intent, Build, Verify & Release, Production and Observe & Improve. A workforce of agents and people spreads across the stages, and a Context & Control layer supplies context, limits and checks to all of them. The system is run against one ratio, validated business value per unit of human intervention.
How is it different from an agentic SDLC?
Agentic SDLC models redraw the familiar lifecycle and place an agent in each phase, which makes them maps of tasks. The AI Engineering Production System is a map of constraints: each stage has a capacity, a queue, an owner and a measure, and the model shows where the constraint sits and where it moves when relieved. It also extends upstream into Direction and the organisational frame, which lifecycle models usually leave out.
What are the stages of the AI Engineering Production System?
Six stages make up the spine. Direction chooses what to work on; Intent turns that into a specification an agent can act on and a test can check; Build produces the change; Verify & Release decides whether it ships; Production lands it; Observe & Improve proves it in use and feeds evidence back to Intent and Direction. Shape, covering enterprise focus, organisation design, funding and decision rights, is the frame around them rather than a stage.
Where do AI agents fit in the engineering operating model?
Agents form part of the workforce across every stage, densest in Build today and spreading stage by stage behind the constraint. Each population of agents needs a scoped mandate, permissions and an escalation path, and draws on a shared layer of context and control. People keep the judgment at every stage, which is where the constraint ends up once the other stages are relieved.
How does MAMOS relate to the AI Engineering Production System?
The spine rewrites the delivery lifecycle in MAMOS's Methods dimension for a workforce that includes agents. MAMOS then serves as the design checklist for each stage, asking what has to change in methods, architecture, management, organisation and skills. Constraints often sit where those dimensions meet, so a redesign that touches only one tends to miss them.
Where should a CTO start redesigning engineering for AI agents?
Start with one high-value flow. Measure the queue time in front of each of the six stages to find the current constraint, then check whether a team boundary, funding line or decision right in the frame is holding it there. Redesign the flow end to end and set a baseline for validated business value per unit of human intervention with a 30-day test.
