Runtime as featured inForbesRead the article

How to Reduce AI Agent Token Costs: A Guide for Finance Teams

Cut AI agent token and compute costs with evals, model routing, open-source models, prompt caching, and deterministic scripts. Worked payments examples.

Gus TrigosCo-founder and CEO, RuntimeUpdated October 4, 20269 min readfinance
ai-agent-costsllm-token-costsmodel-routingopen-source-modelsprompt-cachingfinopscfo

Every agent run has a meter on it. Tokens tick by the thousand, a machine runs by the minute, and at the end of the month the meter shows up as a line on your bill.

Most finance teams see that line for the first time after a pilot, and it's bigger than anyone planned. The good news is that the meter isn't fixed. The same work can cost a fraction of what it did in week one, if someone is actually working to slow the meter down.

That's what we built Runtime to do.

The short answer. You reduce AI agent token costs by measuring cost per run, then pulling four levers. Route each task to the smallest model that passes evals built from your SOPs. Sharpen skills so agents take fewer steps. Cache the parts of every prompt that never change. Turn proven runs into deterministic scripts, so an agent only steps in on edge cases. Keep frontier models as the fallback for the runs that fail or escalate. Runtime does all of this inside your cloud and shows token and compute cost by agent, team, and run.

Read the Meter

You can't cut what you can't see. Before anything else, get cost per run broken out three ways: by agent, by team, and by individual run, with tokens and compute separated.

In Runtime that view is built in. Every run is stored with its steps, the model each step used, the tokens it burned, the minutes its computer ran, and what it cost. That's the number a CFO actually needs. Not the monthly invoice from a model provider, but what it costs to resolve one payout escalation or one merchant re-review, and whether that's going down.

Our agent cost calculator shows how the pieces add up at different volumes and model mixes. The short version: tokens are most of the bill, so model choice matters more than machine size.

Try it with your own volume. Move the sliders to see what routing, scripts, caching, and auto-pause do to cost per run.

Agent cost estimatorrates from runtm.com/agent-costs
All on a frontier model$3.75 / run · $1,876 / mo
With Runtime$0.73 / run · $365 / mo

Who does the work

FrontierOpen modelScript

81%

less per run

$18,129

saved per year

Opus 5.5 and DeepSeek-V3.2 per-minute rates, 8 GB compute. Caching and auto-pause savings are illustrative. Escalated runs re-run half their minutes on the frontier model.

Evals From Your SOPs

The fastest way to overspend is to run everything on the most expensive model because nobody knows if a cheaper one is good enough. Evals answer that question.

We build them from the thing your team already has: the SOP. Each step in the procedure becomes a check. Did the agent find the root cause? Did it cite the evidence? Did it stop at the approval point? Does the reply match policy? Then we grade candidate models against cases your team already resolved, so "good enough" means good enough on your own work.

From there, the loop is simple:

  1. Run the job on a frontier model to set the bar.
  2. Test cheaper models, including open-source ones, against the same checks.
  3. Promote a cheaper model only when it matches the bar on your cases.
  4. Keep the frontier model as the fallback.

That last step is what makes it safe. If a run on the cheaper model fails a check, errors out, or needs judgment it isn't confident about, Runtime escalates the run to a frontier model or better. Anything that moves money still waits for a person. You pay frontier prices only on the runs that need them. We go deeper on evals and routing on the Runtime Lab page.

Open Models, Your Cloud

Open-weight models have closed most of the gap on routine work. Enrichment, data pulls, summaries, and first drafts often pass the same evals on models like Kimi K3 or DeepSeek at a fraction of the price. On our cost model, a minute of DeepSeek-V3.2 costs about an eighth of a minute of Opus 5.5.

There's a second benefit for payment teams. You can serve open models from your own cloud, so sensitive data never leaves your environment.

We keep a running comparison of the best open source models for agents, plus picks by job: payments customer support, merchant underwriting, and transaction monitoring.

Fewer Tokens Per Step

Routing changes the price per token. These change how many tokens you use.

  • Prompt caching. Agents resend the same system prompt, SOP, and tool definitions on every step. Most major providers bill cached input tokens at a fraction of the normal input price. Keep that prefix stable and put the parts that change at the end.
  • Sharper skills. Agents in Runtime learn from resolved cases and propose updates to their own skills. A skill that knows to check sibling MIDs first doesn't spend ten tool calls finding out.
  • Tight context. Send the records a step needs, not the whole account. Truncate long tool outputs to what matters.
  • Batch the work that can wait. Overnight re-reviews don't need real-time answers, and several providers discount batch requests.

Scripts for the Routine Path

This is the biggest lever, and the one most teams miss.

Once an agent has run a job the same way enough times, most of that run is no longer judgment. It's the same five data pulls and the same formatting every time. Paying a model to rediscover those steps on every run is waste.

So in Runtime, engineering can promote a proven run into a reviewed, deterministic script stored in your repo, and owns it like any other code. The script handles the routine path at zero token cost, and it calls the agent only when something doesn't fit: an unusual return code, a mismatch, a pattern it hasn't seen. The agent's tokens go where the judgment is.

Over a few months, the mix shifts. Volume grows, frontier models handle a shrinking share of the work, and scripts take over the routine path.

123$0.42 / run$0.03 / run10,000 runs / weekWeek 1Week 24Frontier modelOpen modelDeterministic scriptCost per run
Illustrative: runs grow 10x while the work shifts from frontier models to open models and scripts. Cost per run falls 93%, and total weekly spend falls 29%.

Compute, Not Just Tokens

Tokens are most of the bill, but compute adds up at volume, especially when agents wait. An agent waiting two hours for an approval shouldn't be billing for two hours of machine time.

Runtime manages that for you:

  • Auto-pause when an agent is idle or waiting on a person, and auto-resume the moment it's needed again.
  • Pre-warming for queues with predictable load, so the first run of the morning doesn't wait on a cold start.
  • Snapshots so a paused computer comes back exactly where it was, without rebuilding its environment.
  • Right-sizing so a summary job doesn't run on the same machine as a code investigation.

We've written more about why persistent runtimes matter and how we think about sandboxes for agents.

Two Worked Examples

Here's what this looks like on two jobs payment teams run every day. Rates come from Runtime's cost model: per minute of agent session, $0.20 on Opus 5.5, $0.108 on Sonnet 5.5, and $0.0237 on DeepSeek-V3.2, with compute at about $0.0084 a minute on 8 GB machines and $0.0168 on 16 GB. Your numbers will differ, and that's what evals and the cost view are for.

A support ticket that becomes an engineering escalation

A merchant says a payout is missing. Support can't see why, so it escalates. The agent has to check transaction monitoring, trace the payout through the ledger and processor, and read the code path in the repo to explain a failed retry. It's a 25-minute, high-complexity run on a 16 GB machine.

All on a frontier modelWith Runtime
Triage and data pulls (12 min)Opus 5.5DeepSeek-V3.2: $0.28
Repo investigation (10 min)Opus 5.5Sonnet 5.5: $1.08
Frontier fallbackEvery run1 in 10 runs: $0.50
Compute$0.42$0.37, paused while waiting on engineering
Cost per run$5.42$2.23 (59% less)
300 escalations a month$1,626$669

The SLA matters as much as the cost. Set it explicitly, like a drafted answer for the merchant within 30 minutes, and use the fallback to protect it: if the cheaper path stalls, the run moves to a frontier model rather than missing the window.

A merchant re-review for underwriting and risk

A merchant's dispute ratio spikes, and risk needs a re-review: KYB, disputes, ownership changes, and sibling MIDs, ending in a reserve recommendation for a person to approve. It's a 17.5-minute run on an 8 GB machine.

All on a frontier modelWith Runtime
Data pullsOpus 5.5, every runDeterministic script: $0.03 compute, no tokens
Recommendation draft (5 min)Opus 5.5DeepSeek-V3.2: $0.16 with compute
Edge casesEvery run15% escalate to Opus: $0.55
Cost per run$3.65$0.73 (80% less)
500 re-reviews a month$1,824$366

The second example saves more because the routine path became a script. That's the pattern: the more a job repeats, the more of it should run without a model at all.

Frequently asked questions

How do you reduce AI agent token costs?

Measure cost per run first, then pull four levers: route each task to the smallest model that passes your evals, sharpen skills so agents take fewer steps, cache the parts of the prompt that never change, and turn proven runs into deterministic scripts that only call an agent on edge cases. Runtime does all four and shows token and compute cost by agent, team, and run.

Are open-source models good enough for financial operations?

For a lot of the work, yes. Alert enrichment, data pulls, summarization, and first drafts often pass the same evals on open-weight models like DeepSeek or Kimi at a fraction of frontier prices. Hard investigations and anything near a decision on money usually stay on frontier models. Evals tell you which is which, task by task.

What happens when a cheaper model gets it wrong?

The run escalates. In Runtime, a run that fails an eval check, errors, or hits low confidence retries on a frontier model, and anything that moves money still waits for a person to approve. You pay frontier prices only for the cases that need them.

How much does prompt caching save?

It depends on how much of each request repeats. Agents resend the same system prompt, SOP, and tool definitions on every step, and most major providers bill cached input tokens at a fraction of the normal input price. Keeping that prefix stable is one of the cheapest savings you can get.

Does compute matter as much as tokens?

Usually less, but it adds up at volume. On Runtime's cost model, a 16 GB agent computer costs about $0.017 per minute, against $0.20 per minute of tokens on a frontier model. Auto-pause, auto-resume, pre-warming, and snapshots keep you from paying for machines that sit idle.

Watch the Meter

You don't need to guess which model to use, or accept the first invoice as the price of agents. Measure cost per run, let evals decide where cheaper models are good enough, keep frontier models as the safety net, and move the routine path into code.

Then watch the meter slow down, run after run.

See your cost per run

Bring one queue and its volume. We'll estimate cost per run on frontier models and with routing, caching, and scripts, inside your own cloud.