Sigma Logic AI Lead with AI. Thrive with Innovation.
Efficiency

Your AI spend is growing and nobody can say what it bought

Inference costs tend to grow faster than usage, and the bill rarely says why. Most of the systems we look at are running a frontier model on work a much cheaper one handles identically, sending the same context hundreds of times a day, and paying for three tools that overlap.

From$4,000

Fixed-fee, read-only: spend attributed per workflow and a ranked reduction list with measured quality risk.

Indicative starting price. The fixed fee for your scope is quoted after the two-day diagnosis, before any build begins.

What makes it work

Three things we insist on

01

Spend attributed to workflows

A provider invoice is one number. We break it down by workflow, feature and user group, so you can see what each thing you built actually costs to run per month and per task.

02

Cheaper only where quality holds

Every proposed reduction is tested against an evaluation set built from your real cases. If moving a workflow to a smaller model costs accuracy, we say so and leave it. Cost work that quietly degrades output is not a saving.

03

The cheapest token is one you never send

Most of the reduction is structural rather than a model swap - caching near-duplicate requests, trimming context that was never read, routing easy cases away from the expensive path, and deleting work nothing consumes.

Attribute the spend first By workflow, model and step. Optimising the part you assumed was expensive is effort wasted
Find which of six causes it is Retrieved context grew, retries are invisible, an agent loop has no ceiling, everything runs on the expensive model, nothing is cached, or a batch job is billing at production rates
Measure quality before changing anything The baseline that makes a cost cut checkable rather than a hope
Change one thing, re-run the set Cost down and quality held, or the change goes back
Every cost change needs an evaluation either side. Cutting context without measuring trades a cost problem for an accuracy problem you will not detect for weeks - and by then nobody remembers what changed.

Capabilities

What is actually included

  1. 01

    Spend attribution

    Provider and infrastructure cost broken down per workflow, with cost per task and month-on-month trend.

  2. 02

    Model right-sizing

    Candidate downgrades tested against your evaluation set, reported with the measured quality delta.

  3. 03

    Caching and routing design

    Identifying near-duplicate traffic and easy cases that never needed the expensive path.

  4. 04

    Context and prompt efficiency

    Finding the context being sent on every call that no output ever depends on.

  5. 05

    Tool and licence overlap

    Mapping what you pay for against what is used, including the seats nobody has opened this quarter.

In detail

What this covers, specifically

A category name is not a scope. These are the individual pieces of work inside this practice - take the two that apply to you and ignore the rest.

  • 01

    Spend attribution

    Provider and infrastructure cost broken down per workflow, feature and user group, with cost per task and a month-on-month trend.

  • 02

    Model right-sizing

    Testing whether a cheaper or smaller model handles a given workflow identically, and reporting the measured quality delta rather than assuming.

  • 03

    Context efficiency

    Finding the context sent on every single call that no output has ever depended on, which is usually the single largest avoidable cost.

  • 04

    Caching strategy

    Identifying near-duplicate requests, because in most production systems a meaningful share of traffic is a question already answered.

  • 05

    Request routing

    Cascading easy cases to a cheap model and escalating only what needs the expensive path, with the escalation rule measured.

  • 06

    Batch versus realtime

    Finding the workloads paying a realtime premium for output nobody reads until the next morning.

  • 07

    Licence and tool overlap

    Mapping what you pay for against what is used, including the seats nobody has opened this quarter and the three tools doing one job.

  • 08

    Infrastructure review

    GPU utilisation, instance sizing and idle capacity for anything self-hosted.

  • 09

    Budget alerting

    Forecasting and thresholds so a runaway loop or a pricing change is caught in hours, not on the invoice.

What you receive

Concrete artefacts, not a slide deck

  • Cost breakdown per workflow, feature and user group
  • Ranked reduction list with measured quality risk against each
  • Evaluation results for every proposed model change
  • Caching, routing and context recommendations
  • Implementation plan, if you want us to make the changes
01 Diagnose Days 1-2
02 Prove Week 1
03 Integrate Weeks 2-3
04 Operate Ongoing

Around three weeks end to end. That comes from scoping tightly to one workflow - not from skipping a phase. Each still ends in evidence you can check.

Let's talk

Fixed fee, read-only, one week

We need billing exports and read access to the systems. We change nothing. You get a written breakdown of where the money goes and a ranked list of reductions with the quality risk stated against each - yours to implement with us or without us.