Spend attributed to workflows
A provider invoice is one number. We break it down by workflow, feature and user group, so you can see what each thing you built actually costs to run per month and per task.
Inference costs tend to grow faster than usage, and the bill rarely says why. Most of the systems we look at are running a frontier model on work a much cheaper one handles identically, sending the same context hundreds of times a day, and paying for three tools that overlap.
From$4,000
Fixed-fee, read-only: spend attributed per workflow and a ranked reduction list with measured quality risk.
Indicative starting price. The fixed fee for your scope is quoted after the two-day diagnosis, before any build begins.
What makes it work
A provider invoice is one number. We break it down by workflow, feature and user group, so you can see what each thing you built actually costs to run per month and per task.
Every proposed reduction is tested against an evaluation set built from your real cases. If moving a workflow to a smaller model costs accuracy, we say so and leave it. Cost work that quietly degrades output is not a saving.
Most of the reduction is structural rather than a model swap - caching near-duplicate requests, trimming context that was never read, routing easy cases away from the expensive path, and deleting work nothing consumes.
Capabilities
Provider and infrastructure cost broken down per workflow, with cost per task and month-on-month trend.
Candidate downgrades tested against your evaluation set, reported with the measured quality delta.
Identifying near-duplicate traffic and easy cases that never needed the expensive path.
Finding the context being sent on every call that no output ever depends on.
Mapping what you pay for against what is used, including the seats nobody has opened this quarter.
In detail
A category name is not a scope. These are the individual pieces of work inside this practice - take the two that apply to you and ignore the rest.
Provider and infrastructure cost broken down per workflow, feature and user group, with cost per task and a month-on-month trend.
Testing whether a cheaper or smaller model handles a given workflow identically, and reporting the measured quality delta rather than assuming.
Finding the context sent on every single call that no output has ever depended on, which is usually the single largest avoidable cost.
Identifying near-duplicate requests, because in most production systems a meaningful share of traffic is a question already answered.
Cascading easy cases to a cheap model and escalating only what needs the expensive path, with the escalation rule measured.
Finding the workloads paying a realtime premium for output nobody reads until the next morning.
Mapping what you pay for against what is used, including the seats nobody has opened this quarter and the three tools doing one job.
GPU utilisation, instance sizing and idle capacity for anything self-hosted.
Forecasting and thresholds so a runaway loop or a pricing change is caught in hours, not on the invoice.
What you receive
Around three weeks end to end. That comes from scoping tightly to one workflow - not from skipping a phase. Each still ends in evidence you can check.
Further reading
Let's talk
We need billing exports and read access to the systems. We change nothing. You get a written breakdown of where the money goes and a ranked list of reductions with the quality risk stated against each - yours to implement with us or without us.
Related practices