Sigma Logic AI Lead with AI. Thrive with Innovation.
Assurance

Prove it works before your customers test it for you

"It seems good" is the only quality signal most teams have. It is also the signal that quietly degrades when a provider updates a model, your product changes, or the questions people ask move on. Evaluation replaces the impression with a number you can act on.

From$8,000

Evaluation set from your real cases, scoring harness, measured baseline and CI integration.

Indicative starting price. The fixed fee for your scope is quoted after the two-day diagnosis, before any build begins.

What makes it work

Three things we insist on

01

Built from your cases, not a benchmark

Public benchmark scores tell you nothing about whether a system handles your customers, your terminology and your edge cases. The set we build comes out of your own history, including the awkward examples people remember.

02

Every change re-measured

Evaluations run in your pipeline on each release, so a prompt edit or a model change produces a reported delta instead of a hopeful deployment. Quality stops being a matter of opinion in a standup.

03

We try to break it

Ambiguous inputs, contradictory instructions, out-of-scope requests, adversarial phrasing and the long tail nobody wrote a test for. Better we find the failure than a customer, or a screenshot on social media.

Your resolved tickets Real cases that already happened, not invented ones
Labelled evaluation set Stratified by category, including the cases the system should refuse
Every change runs against it Prompt edits, model swaps, retrieval changes - scored per category, never on the average
Compared to the last version Not to a benchmark. To what your own system did yesterday
Clears the bar Ships, with the number recorded so the next change has something to beat Release
Regression Caught here rather than by a customer three weeks later Blocked
The gate is the product, not the set. A set nobody runs on every change is a document. Production failures feed back in as new cases, so it grows exactly where the system actually fails.

Capabilities

What is actually included

  1. 01

    Evaluation set construction

    A held-out set drawn from your real cases, labelled with the correct outcome and weighted toward what actually matters.

  2. 02

    Scoring harness

    Automated scoring with model-as-judge where appropriate, calibrated against human ratings so the score means something.

  3. 03

    Regression testing in CI

    Wired into your pipeline so every change is measured before it ships, not after.

  4. 04

    Robustness probing

    Structured attempts to induce wrong, unsafe or out-of-scope behaviour through input alone. This is behavioural testing, not a security assessment - formal penetration testing is a separate discipline and we will refer you.

  5. 05

    Per-segment reporting

    Accuracy broken down by language, customer type, channel or region, because a strong average routinely hides a weak segment.

  6. 06

    Acceptance criteria

    An agreed bar for go and no-go, so launch decisions stop being a judgement call under deadline pressure.

In detail

What this covers, specifically

A category name is not a scope. These are the individual pieces of work inside this practice - take the two that apply to you and ignore the rest.

  • 01

    Evaluation set construction

    A held-out set drawn from your real historical cases, labelled with the correct outcome and weighted toward what actually matters commercially.

  • 02

    Rubric design

    Defining what "correct" means for your task, which is harder and more valuable than it sounds when outputs are free text rather than a label.

  • 03

    Automated scoring harness

    The code that runs the set and produces a score, using model-as-judge where appropriate and calibrated against human ratings so the number is trustworthy.

  • 04

    Human review and annotation

    A workflow for the cases automated scoring cannot settle, with inter-rater agreement measured rather than assumed.

  • 05

    Regression testing in CI

    Evaluations wired into your pipeline so every prompt edit, model change and dependency upgrade produces a reported delta before it ships.

  • 06

    Robustness probing

    Structured attempts to induce wrong, unsafe or out-of-scope behaviour through input alone. Behavioural testing, not a security assessment; formal penetration testing is a separate discipline.

  • 07

    Per-segment reporting

    Accuracy broken down by language, customer type, channel or region, because a healthy average routinely conceals one badly served group.

  • 08

    Confidence calibration

    Checking that when the system reports 80% confidence it is right about 80% of the time, without which any routing threshold is arbitrary.

  • 09

    Provider comparison

    Running the same evaluation across candidate models so a switch is decided on measured quality and cost rather than a launch announcement.

  • 10

    Acceptance criteria

    A written bar for go and no-go, agreed in advance, so launch decisions are not made on instinct under deadline pressure.

What you receive

Concrete artefacts, not a slide deck

  • Evaluation set built from your real cases, owned by you
  • Scoring harness and rubric, documented
  • Measured baseline with per-segment breakdown
  • CI integration so every change is tested automatically
  • Written acceptance criteria for launch decisions
01 Diagnose Days 1-2
02 Prove Week 1
03 Integrate Weeks 2-3
04 Operate Ongoing

Around three weeks end to end. That comes from scoping tightly to one workflow - not from skipping a phase. Each still ends in evidence you can check.

Let's talk

Get a baseline on what you already run

Point us at a system that is live and we will build an evaluation set from your real cases and tell you how it currently performs. Most teams have never had that number. It is usually the most useful thing they learn all quarter.