Sigma Logic AI Lead with AI. Thrive with Innovation.
Engineering

When nothing off the shelf understands your business

Some problems are specific to you: your terminology, your rules, your data, your edge cases. We build those systems from scratch - properly evaluated, integrated with what you run, and handed over with the source and the documentation so you are never locked in.

From$20,000

A bespoke system in your environment with an evaluation suite, measured baseline and full handover.

Indicative starting price. The fixed fee for your scope is quoted after the two-day diagnosis, before any build begins.

What makes it work

Three things we insist on

01

Scoped against a decision, not a demo

We start from the decision the system must make and the cost of getting it wrong. That determines the accuracy bar, the architecture and the budget - before anyone writes a prompt.

02

Evaluated before it is trusted

Every build ships with a held-out evaluation set drawn from your real cases, and a measured baseline. You know what the system gets right, what it gets wrong, and how that changes with each release.

03

Built to be handed over

Source, prompts, training data, evaluation sets and runbooks are yours. We would rather be retained because the work is good than because you cannot operate without us.

Scope one workflow, end to end An input, a defined done, a defined escalation, and a number that moves. Not a capability
Decide what it may do Drafts a reply or sends it. Recommends a refund or issues it. This drives the architecture and the audit logging you will wish you had
Build against your systems Real APIs, real permissions, structured output validated strictly against your own schema
Hand over everything Source, prompts, evaluation sets, runbooks. What you own is what you can change without us
The permission decision is the architecture decision. Every step up the ladder multiplies the value and the blast radius, which is why it is settled before anything is built rather than discovered afterwards.

Capabilities

What is actually included

  1. 01

    Retrieval over your knowledge

    Search and question-answering across your documents, tickets, code and records, with citations and access control preserved.

  2. 02

    Fine-tuning and open-weight deployment

    Smaller specialised models where volume, latency, cost or data residency rule out a hosted API.

  3. 03

    Multi-step agents with real tools

    Systems that plan, call your APIs, check their own work and stop when they are unsure.

  4. 04

    Classical ML where it wins

    Forecasting, classification, ranking and anomaly detection - often cheaper and more accurate than a language model for the same job.

  5. 05

    Production engineering

    Versioning, observability, cost controls, rollback and load testing. The unglamorous half that decides whether it survives contact with users.

In detail

What this covers, specifically

A category name is not a scope. These are the individual pieces of work inside this practice - take the two that apply to you and ignore the rest.

  • 01

    Retrieval-augmented generation (RAG)

    Answering from your own documents and records by retrieving the relevant passages at question time and citing them, rather than relying on what a model memorised.

  • 02

    Fine-tuning

    Adapting a smaller model to your domain, terminology and output format using LoRA or similar, where prompting alone cannot reach the accuracy bar.

  • 03

    Open-weight deployment

    Running models inside your own cloud or on-premise, for cases where data residency, contractual terms or volume economics rule out a hosted API.

  • 04

    Multi-step agents

    Systems that plan a task, call your APIs to carry it out, check their own work and stop when uncertain, rather than answering in one shot.

  • 05

    Classical machine learning

    Forecasting, classification, ranking and anomaly detection using conventional models, which are frequently cheaper, faster and more accurate than a language model for the same job.

  • 06

    Structured extraction

    Turning free text, email, documents or transcripts into typed, validated records that another system can consume.

  • 07

    Evaluation harness

    A held-out test set built from your real cases plus the scoring code, so every future change can be measured instead of eyeballed.

  • 08

    Model routing and fallback

    Sending each request to the cheapest model that handles it well, with defined behaviour when a provider is slow, rate-limited or down.

  • 09

    Integration middleware

    The API layer, queues, retries and idempotency that sit between the model and your systems and decide whether it survives production.

What you receive

Concrete artefacts, not a slide deck

  • Technical design with architecture and cost model
  • Working system in your environment
  • Evaluation suite and measured baseline
  • Full source, prompts and documentation
  • Handover and enablement for your team
01 Diagnose Days 1-2
02 Prove Week 1
03 Integrate Weeks 2-3
04 Operate Ongoing

Around three weeks end to end. That comes from scoping tightly to one workflow - not from skipping a phase. Each still ends in evidence you can check.

Let's talk

Bring us the problem, not the spec

The most useful first conversation is about the decision you need made better and what it currently costs you to get it wrong. We will tell you whether AI is genuinely the right tool - and what we would build if it is.