The build

Pricing guides published by development firms are quoted in dollars and converge on a consistent ladder:

  • Prototype: $10,000–$30,000. One workflow, one data source, no production controls.
  • MVP: $20,000–$60,000. Real users, limited scope, manual oversight.
  • Single production agent: $20,000–$80,000 for something straightforward.
  • Mid-market deployment: $40,000–$150,000, which is where most first serious projects land.
  • Enterprise multi-agent system: $400,000 and up, once compliance, custom integrations and an orchestration layer are involved.

The interesting part of that ladder is not the top. It is the jump from MVP to production, where scope stops being about the model and starts being about your systems.

Why integration eats the budget

For most enterprise deployments, integration engineering and QA or safety testing together account for 40–60% of total build cost. That surprises people who assumed the expensive part was the AI.

It should not. A support agent that answers questions is a weekend project. A support agent that reads the order system, checks the refund policy, writes to the CRM, respects a customer's marketing preferences and can explain its decision to an auditor is systems work with a language model attached. The model is the cheapest component in that sentence.

What it costs to run

A production agent serving real users costs roughly $3,200–$13,000 a month once you count model API usage, infrastructure, monitoring, periodic tuning and security maintenance. Annual operating cost typically lands at 15–25% of the initial build. Over three years, total cost of ownership usually reaches one and a half to two times the build price.

Run that arithmetic before signing anything. A mid-market build at the bottom of its range, carrying a run cost in the middle of this one, roughly doubles inside the first year, and most business cases are written against the build number alone.

The line items nobody quotes

Four costs turn up after go-live and appear in almost no proposal.

  • Evaluation maintenance. Every new failure mode becomes a test case. Someone writes and maintains that suite, or quality drifts unnoticed.
  • Model version churn. Providers deprecate and replace models on their own schedule. Each move means re-running evaluations and sometimes re-tuning prompts.
  • Data cleanup. Agents expose how bad the underlying records are. That work is usually discovered, not planned.
  • Human review. Most production agents need a person handling the cases they escalate. That is a staffing line, not a software line.

ONS figures give a hint of how unprepared most organisations are for this. The average UK business using AI runs just 1.6 AI technologies, barely up from 1.4 in late 2023, even though adoption roughly tripled over the same period. Most firms have one shallow deployment and no operational practice around it. The second agent is where that gap turns into a cost.

UK day rates set the floor

This is where a British buyer should stop reading international price guides and look at local labour. On ITJobsWatch contractor data, the median AI engineer rate sits near £550 a day, AI software developers around £610, and lead data scientists around £725. General software contractors sit at roughly £350–£450. London and the South East run 10–20% above the national average, and rates across the wider contractor market span about £200 a day at junior level to £1,200 at principal.

Four people for three months at AI-engineer rates is roughly £150,000 of labour before anyone writes a line of integration code. Any UK quote well below that is one of three things: a much smaller scope than you think, a much more junior team than you were shown, or a platform licence with a services wrapper.

What a defensible quote looks like

A quote you can actually govern separates four things: discovery, build, integration, and the first six months of running the agent. It names who owns the prompts and evaluation suite at the end. It states an expected cost per task at your projected volume, with the arithmetic shown. And it says what happens to the price if that volume triples.

A single blended figure for "AI agent development" hides every one of those, which is convenient for whoever wrote it.

The cheapest decision you can make

Scope the first agent around a workflow that already has a measurable cost. Number of tickets, hours of reconciliation, days of processing delay. If nobody can state that number today, the project has no denominator, and the 380% overrun figure stops being a statistic about other companies.