Scoped against a decision, not a demo
We start from the decision the system must make and the cost of getting it wrong. That determines the accuracy bar, the architecture and the budget - before anyone writes a prompt.
Some problems are specific to you: your terminology, your rules, your data, your edge cases. We build those systems from scratch - properly evaluated, integrated with what you run, and handed over with the source and the documentation so you are never locked in.
From$20,000
A bespoke system in your environment with an evaluation suite, measured baseline and full handover.
Indicative starting price. The fixed fee for your scope is quoted after the two-day diagnosis, before any build begins.
What makes it work
We start from the decision the system must make and the cost of getting it wrong. That determines the accuracy bar, the architecture and the budget - before anyone writes a prompt.
Every build ships with a held-out evaluation set drawn from your real cases, and a measured baseline. You know what the system gets right, what it gets wrong, and how that changes with each release.
Source, prompts, training data, evaluation sets and runbooks are yours. We would rather be retained because the work is good than because you cannot operate without us.
Capabilities
Search and question-answering across your documents, tickets, code and records, with citations and access control preserved.
Smaller specialised models where volume, latency, cost or data residency rule out a hosted API.
Systems that plan, call your APIs, check their own work and stop when they are unsure.
Forecasting, classification, ranking and anomaly detection - often cheaper and more accurate than a language model for the same job.
Versioning, observability, cost controls, rollback and load testing. The unglamorous half that decides whether it survives contact with users.
In detail
A category name is not a scope. These are the individual pieces of work inside this practice - take the two that apply to you and ignore the rest.
Answering from your own documents and records by retrieving the relevant passages at question time and citing them, rather than relying on what a model memorised.
Adapting a smaller model to your domain, terminology and output format using LoRA or similar, where prompting alone cannot reach the accuracy bar.
Running models inside your own cloud or on-premise, for cases where data residency, contractual terms or volume economics rule out a hosted API.
Systems that plan a task, call your APIs to carry it out, check their own work and stop when uncertain, rather than answering in one shot.
Forecasting, classification, ranking and anomaly detection using conventional models, which are frequently cheaper, faster and more accurate than a language model for the same job.
Turning free text, email, documents or transcripts into typed, validated records that another system can consume.
A held-out test set built from your real cases plus the scoring code, so every future change can be measured instead of eyeballed.
Sending each request to the cheapest model that handles it well, with defined behaviour when a provider is slow, rate-limited or down.
The API layer, queues, retries and idempotency that sit between the model and your systems and decide whether it survives production.
What you receive
Around three weeks end to end. That comes from scoping tightly to one workflow - not from skipping a phase. Each still ends in evidence you can check.
Further reading
Let's talk
The most useful first conversation is about the decision you need made better and what it currently costs you to get it wrong. We will tell you whether AI is genuinely the right tool - and what we would build if it is.
Related practices