← Insights · August 18, 2026

What running AI in production actually costs

Article card on a dark navy background reading The build got a budget. The run got a shrug, with the last line highlighted in green.

Every AI business case has two columns. The build column gets weeks of scrutiny, a contingency line and a steering group. The run column gets a licence fee and a shrug.

That shrug is where production AI goes to get expensive. Not dramatically, in one bad invoice, but quietly, in engineering time nobody allocated, guidance nobody was assigned to read and model retirements nobody saw coming.

We build and run AI systems for a living, so treat what follows as a practitioner’s account with a declared interest. The costs are real either way.

The spend is now mainstream. The discipline isn’t.

The State of FinOps 2026 survey, the FinOps Foundation’s sixth annual survey of its practitioner community, published February 2026 with 1,192 respondents representing more than $83bn in annual cloud spend, found that 98% of teams now manage AI spend. Two years ago it was 31%. (The Foundation is a Linux Foundation programme and does not publish the fieldwork window on the report page, so read it as a late 2025 snapshot of a practitioner community. The Linux Foundation’s own release carries the same figures.)

Three details in that report matter more than the headline.

AI cost management is the single most wanted skillset across every organisation size. The top tooling request, ahead of everything else, is granular monitoring of AI spend: tokens, LLM requests and GPU utilisation. The report’s most honest line comes from a practitioner: “Is your AI providing value? No one can answer that question yet.”

Nearly everyone is now paying to run AI. Almost nobody can yet see what they are getting for it. That gap is the run column, and it has five entries the business case usually leaves out.

Cost one: the model underneath you retires on someone else’s schedule

Buy a boiler and it is yours for fifteen years. Build on a hosted model and the vendor decides how long “done” lasts.

OpenAI’s own deprecations page is worth ten minutes of any CTO’s time (all dates below checked against it on 10 August 2026). The Assistants API is removed on 26 August 2026, a year after notice was given. DALL·E 2 and 3 came out of the API on 12 May 2026. A widely used GPT-4o snapshot retires on 23 October 2026. Preview models can go with as little as two weeks’ notice. Every entry has a recommended replacement, which is helpful, and every replacement behaves slightly differently, which is the point.

A model swap is not a find and replace. The prompts that worked need re-testing. The outputs your downstream process depends on need re-validating. The edge cases you fixed in March need re-fixing, or at least re-checking, against a model with different habits. If nobody budgeted for that work, it comes out of whatever your engineers were supposed to be doing instead.

None of this is a criticism of OpenAI, who give more notice than most. It is a criticism of business cases that treat a hosted model like a boiler.

Cost two: evaluation is a running cost, not a launch task

Most teams evaluate once, at go-live, against a test set someone built in a hurry. Then the system changes underneath them: a model upgrade, a prompt tweak, a new document format, drift in what users actually ask.

The unglamorous fix is an evaluation suite that runs before and after every change, with a threshold that can block a release. Building it is a one-off cost. Running it, keeping the test set honest and arguing about the threshold is forever. It is the AI equivalent of regression testing, and skipping it has the same failure mode: everything works until the week it doesn’t, and nobody can say which change did it.

Cost three: someone has to be watching

McKinsey’s 2026 AI trust survey (around 500 organisations, respondents responsible for AI governance and risk, fieldwork December 2025 to January 2026) found roughly 8% reporting AI incidents, and of those affected, almost 60% rated their own response satisfactory or worse. Nearly two-thirds named security and risk, not regulation and not technical limits, as the top barrier to scaling agentic AI.

Retool’s State of AI Governance 2026 (a vendor survey, 307 CTOs, CIOs and CISOs, published June 2026) found only 5% very confident they have full visibility of what is running in production. It also found 51% unable to confirm whether AI-generated tools had caused incidents.

Read together: incidents happen, responses disappoint and half the leadership can’t see the estate. Monitoring is not a dashboard you buy. It is a definition of what “drift” and “failure” mean numerically for your system, plus a named person who gets paged, plus the hours they spend responding. All three recur monthly.

Cost four: the regulatory reading list keeps arriving

Checked against the ICO’s guidance pipeline on 10 August 2026: final automated decision-making guidance, previously expected this summer, is now due Winter 2026. Agentic AI guidance is also due Winter 2026. Foundation model guidance is due Summer 2026. In financial services, the Treasury Committee said in January that regulators’ wait-and-see approach is not doing enough and asked the FCA to publish senior manager accountability guidance by the end of 2026.

Every one of those documents, when it lands, has to be read, mapped against your deployed systems and turned into changes. If your AI screening tool touches decisions about people, the winter ADM guidance is not optional reading. The NCSC’s secure AI guidelines already devote a quarter of their structure to secure operation and maintenance, which tells you what the UK’s security establishment thinks the long half of the job is.

Absorbing guidance is a running cost. It arrives on the regulator’s schedule, like model retirements arrive on the vendor’s. Your budget is the only place your own schedule exists.

Cost five: a name in the box

Every cost above is survivable if one question has an answer: whose job is this?

Someone owns the approved model list and hears about deprecations before the deadline is a fortnight out. Someone owns the evaluation thresholds. Someone gets paged. Someone reads the ICO so the rest of the organisation doesn’t have to. In a bank this is recognisable as model risk management and it is already mandatory. Everywhere else it is optional, which is why in most organisations it is nobody, and the run column becomes unfunded overtime plus luck.

The person does not need to be full time. The name needs to exist.

What this does to the business case

Put the run column in before the pilot, not after. In our own method (APEX), Monitor is a phase with the same standing as Design and Implement, because a system nobody is resourced to run is not an asset. It is a liability with good first impressions.

Two honest caveats. First, per-token prices keep falling, so the raw compute line often shrinks. The costs that don’t shrink are the human ones: evaluation, monitoring, migration and reading. Second, none of this is an argument against running AI in production. We do it, our clients do it and the returns are real when the system keeps working. It is an argument against pretending the keeping-working part is free.

The narrow version of this argument is also the cheapest to act on: give the system the minimum autonomy the job needs. Every notch of autonomy you don’t grant is monitoring you don’t have to build and incidents you don’t have to respond to.

Five things to do this month

Ask what your AI estate actually is; if the answer takes more than a day, that is finding one. Ask which vendor models it depends on and check their deprecation pages. Ask what gets re-tested when a model changes and who decided the threshold. Ask who is named for monitoring and for regulatory guidance, and what those hours cost. Then write the three-year run column with those numbers in it, and see if the business case still smiles.

If it does, you have a project. If it doesn’t, you have just saved yourself the expensive way of finding out.

Occasional, useful notes on applied AI, including what the new regulation means for UK businesses: subscribe to Insights on aiapplied.uk (form in the footer). No spam.

Insights

Occasional, useful notes on applied AI.

What's actually working, what to ignore, and what the new regulation means for UK businesses. No spam.

We’ll only use your email address to send you these Insights notes. We never share it, and you can unsubscribe from any email. See our Privacy Policy.

AI Services

Where we work

Company

Latest writing

AI Applied Ltd, Technology House, 9 Newton Place, Glasgow G3 7PR. Registered in Scotland SC806963. support@aiapplied.uk · +44 141 465 5233