Skip to content

Perspectives

Delivering AI systems

Why conventional delivery methods do not transfer cleanly — and what a delivery framework for AI has to account for.

· ai, delivery, frameworks

Saviour Alfino · 3 October 2026

Why AI projects need a different kind of delivery

Most organisations starting with AI do not fail because the model is weak; they fail because the project around it was never designed for uncertainty. A conventional software project can usually promise a date, because the team knows the thing can be built. An AI project often cannot, because nobody yet knows whether the system will be accurate enough, cheap enough or trusted enough to use.

That difference changes what good delivery looks like. Quality is measured statistically rather than demonstrated once; behaviour can shift when a provider updates a model; and adoption is never automatic, since people resist systems they do not understand or were oversold.

I have put together a delivery framework that keeps the familiar shape of stage-gated project management whilst embedding the AI-specific decisions beneath it. I call it AI-Systems Development Method (inspired from DSDM). It is not a new methodology to learn; it is conventional delivery, with the AI depth placed where it belongs. This post walks through it for organisations taking their first serious steps.

The framework at a glance

AI-SDM has six phases, each ending in a gate, and six workstreams that run across all of them. Each phase answers one plain question, and each gate asks for evidence rather than opinion.

PhaseThe question it answersGate passes when
1. Initiate & feasibilityShould we build this at all?The outcome is credible, the data feasible, the risk acceptable and success measurable
2. Analyse & designWhat are we building this time?We know what we are building, how it works, what it costs and who owns it
3. Build & experimentCan we make it perform?The build is reproducible, an evaluation baseline exists and regressions are blocked
4. Integrate & hardenDoes it hold in the real organisation?It works under realistic users, data, load and attack
5. Assure & releaseIs it safe and good enough to release?Independent assurance and the business both say yes
6. Deploy & operateCan we operate and control it?Ownership, monitoring, economics, rollback and continuous evaluation are in place
The AI System Development Method: an initiate-and-feasibility phase, then proof of concept, MVP and scaled solution tracks through five further phases with gates, supported by project controls, security and guardrails, adoption, and a continuous feedback loop.
The AI System Development Method at a glance

The six workstreams are AI and model; data and knowledge; application and agent; trust and security; quality and operations; and platform and infrastructure. Each has a named owner from the design phase onwards, so no part of the system is left without someone accountable for it.

Three release tracks, and gates that bend without breaking

Not every AI initiative needs the full journey on day one. AI-SDM runs the same phases at three levels of ambition: a proof of concept, a minimum viable product, and a scaled solution.

  • Proof of concept answers whether the idea works at all. It can skip the hardening phase, because nobody should rely on it yet.
  • Minimum viable product puts the system in front of real users. Hardening becomes mandatory here; a proof of concept promoted straight into production is a common and costly mistake.
  • Scaled solution extends a proven system to more users, data and use cases, with full assurance at every release.

The gates themselves are deliberately not pass-or-fail. Each control can be mandatory, conditional, deferred, accepted as a named risk, or not applicable to this track. A gate that only says yes or no turns governance into bureaucracy, and teaches teams to route around it.

The same honesty applies to planning. Feasibility experiments in the first phase are time-boxed, and each has a kill criterion agreed before it starts. Dates beyond the design gate stay indicative, because committing to a date before you know something is possible is a promise nobody can keep.

The six phases in practice

Every phase follows the same pattern: what the project manager manages, which AI decisions must be made, and what evidence the gate expects. The project management will feel familiar; the AI decisions are where most organisations need help.

1. Initiate & feasibility. The first decision is whether you need generative AI at all; conventional software or classical machine learning is sometimes better and cheaper. Decide how much autonomy the system needs (assist, automate or orchestrate) and what a wrong answer would cost. Build the business case on cost per task at expected and stretch volumes, and confirm the data is licensed for this use.

2. Analyse & design. Choose the model route, with a named fallback; how the system gets its knowledge, whether through the prompt, retrieval, adaptation or a mix; and the pattern, from a single call through a fixed workflow to a bounded agent. Quantify the non-functional requirements, such as accuracy, latency, availability and cost. Record each decision with its alternatives, rationale, consequence and owner.

3. Build & experiment. Build the application, the AI capability and the evaluation together, not in sequence. Treat prompts as engineering artefacts under version control. Measure retrieval separately from answer quality, and let the pipeline block any change that lowers the evaluation score.

4. Integrate & harden. Prove that retrieval respects source permissions, using a restricted account. Set ceilings on steps, time and spend for anything agent-like, and decide whether failures degrade, fall back or stop. Run a pilot with exit criteria fixed in advance, alongside load, failover and adversarial testing.

5. Assure & release. Prove quality statistically, with sample size and confidence stated, rather than through a demonstration. Evaluate the whole system, not model benchmarks. Where automated scoring is used, calibrate it against human graders. Assurance should sit outside the build team.

6. Deploy & operate. Hand the system to a funded business owner, and track benefits against the business case. Re-test the live system on a schedule, because a provider can change the model underneath you. Make sure the model, prompts, index and safety rules can each be rolled back independently. Go-live is not closure.

Three threads that run from start to finish

Some concerns cannot be handed to a single phase. AI-SDM treats three of them as continuous threads.

Security. Security principles are set in the first phase, threat-modelled in design, tested in hardening and monitored in operation. The OWASP Top 10 for LLM and GenAI is a useful shared checklist; its 2025 edition covers risks such as prompt injection, sensitive information disclosure, excessive agency and unbounded consumption.

Evaluation. Evaluation is built alongside the system from the start, not bolted on before release. Many teams use one language model to grade another, which scales well but carries known biases. Zheng et al. (2023) found strong position bias in every judge model they tested, along with a tendency to favour longer answers; they also examined whether judges prefer their own outputs. Calibrating automated scores against human graders guards against all three.

Adoption. A system that works can still fail if people do not use it. Resistance, fear for jobs and disappointment after over-promising are delivery risks, not afterthoughts. The mitigations are ordinary but often skipped: honest communication without hype, training, governance from day one, and a pilot that shows measurable impact before wider rollout.

Where to start

If your organisation is about to begin its first AI project, five habits will matter more than any tool choice.

  1. Ask whether you need generative AI at all before you commit budget to it.
  2. Agree success measures with numbers attached, and a kill criterion for every experiment.
  3. Build evaluation from the first sprint, and keep a held-out test set nobody tunes against.
  4. Never promote a proof of concept to real users without hardening it first.
  5. Fund an owner for the system after go-live, and keep testing it once it is live.

None of this requires a new methodology. It requires the discipline of good project delivery, applied with an understanding of where AI behaves differently. That combination is what turns a promising demonstration into a system people use and trust.

Sources and further reading