AI & Machine Learning
Production AI engineering for ambitious teams — LLM agents, RAG pipelines, and bespoke ML models that survive the trip from notebook to customer.
01 — Approach
How we engage.
- LLM-powered agents & copilots — multi-step agents on LangGraph, structured tool-use, planning, and memory architectures that don't hallucinate critical fields.
- Retrieval-augmented generation (RAG) pipelines — vector + lexical hybrid retrieval, query rewriting, reranking, and grounded answer generation over your knowledge base, docs, or product catalogue.
- Custom AI integrations — OpenAI, Anthropic Claude, Google Gemini, Mistral, and open-weights models served on Modal, Replicate, or your own GPUs.
- Fine-tuning & distillation — LoRA / QLoRA on Llama, Mistral, and Qwen; distilling expensive frontier-model traces into cheaper production models you can actually afford to run.
- Predictive ML models — churn, fraud, demand forecasting, ranking, and recommendation systems built on PyTorch, XGBoost, and the boring tools that still win.
- AI evaluation & observability platforms — eval harnesses, regression suites, prompt versioning, and the dashboards your team needs to ship LLM features without flying blind.
- Production-first. Every system ships with evals, regression tests, observability, and a written runbook. The notebook is the start, not the deliverable.
- Senior-only delivery. A staff-level ML engineer owns the project from kickoff to launch — most of our team have shipped production AI at startups and large platforms before joining.
- Honest about what AI can and cannot do. We turn down briefs that need more determinism than an LLM can give. You will hear "no" as often as "yes".
- Cost and latency budgets. Token spend, p95 latency, and per-request cost are tracked from day one — not "optimised in phase two".
- Model-agnostic. We pick the right model for the job — frontier when it earns its keep, open-weights when it doesn't. Vendor lock-in is a smell.
- Safety & guardrails built in. Prompt-injection defenses, PII redaction, content filters, and red-team playbooks shipped as part of every customer-facing system.
- Frontier LLMs: Anthropic Claude (Opus / Sonnet / Haiku), OpenAI GPT-4 / o-series, Google Gemini, Mistral Large.
- Open-weights: Llama 3 / 4, Mistral 7B–8x22B, Qwen 2.5, served on Modal, vLLM, or your own H100s.
- Agents & orchestration: LangGraph, LangChain, custom-built when the framework gets in the way.
- RAG & vector search: pgvector, Pinecone, Weaviate, Qdrant, BM25 hybrid, Cohere & Voyage rerankers.
- Fine-tuning & ML: Hugging Face Transformers, PyTorch, Axolotl, Unsloth, Modal for GPU jobs.
- Eval & observability: Braintrust, LangSmith, in-house eval harnesses, Posthog for product analytics on AI features.
- Backend & infra: Python (FastAPI), Node.js, Rust where latency demands it; Kubernetes on AWS or GCP.
- Fixed-scope build — $75k–$400k+, two to six months. Best for new AI products with a clear user task.
- AI feature pilot — $40k–$100k, four to eight weeks. Best for bolting a focused AI capability onto an existing product.
- Embedded squad — monthly retainer, two to four engineers. Best for ongoing AI evolution alongside your internal team.
- Discovery sprint — $25k, two weeks. A de-risking exercise that produces an eval set, an architecture, and a fixed quote.
02 — What's included
Every engagement ships with.
Senior lead
A 10+-year practitioner who stays on the work, end-to-end.
Design system
A scalable foundation, not screen-by-screen one-offs.
Production deploys
Fortnightly increments to a staging URL.
Documentation
Runbooks, ADRs, and onboarding materials.
03 — Process
Four phases. Always.
Discovery
1–2 weeks. Audit, listen, scope.
Design
2–4 weeks. Prototypes you can click.
Build
6–16 weeks. Two-week cadences.
Stewardship
Ongoing. Continuity beats handoff.
04 — Common questions
Frequently Asked Questions
How much does AI development cost?
Most of our bespoke AI engagements land between $75,000 and $400,000 for a fixed-scope build. Smaller AI feature pilots start at around $40,000, and two-week discovery sprints (eval set + architecture + fixed quote) are $25,000. We publish honest ranges because we would rather discuss budget on the first call than weeks into a proposal cycle.
How long does it take to ship an AI feature to production?
A focused AI feature pilot ships in 4–8 weeks. A larger AI product or agent system runs 12–24 weeks end-to-end, including eval harness, safety review, and a hardened launch. We work in two-week sprints and measure every change against the eval set, so progress is visible fortnightly — not at month three.
Should we use Claude, GPT-4, Gemini, or an open-weights model?
We are model-agnostic and will tell you honestly. Frontier models (Claude, GPT-4, Gemini) are the right call when reasoning quality is the bottleneck and per-call cost is a rounding error against the value. Open-weights (Llama, Mistral, Qwen) are the right call at high volume, for on-prem compliance, or when fine-tuning beats frontier on a narrow domain. Most engagements are frontier-first with a distilled fallback for the tail.
Do you build RAG pipelines, or only LLM integrations?
RAG is a core practice. We build hybrid retrieval (vector + lexical), query rewriting, reranking, and grounded answer generation over your knowledge bases, documents, product catalogues, or transactional data. We have shipped production RAG for legal, healthcare, fintech, and hospitality clients.
Can you fine-tune a model on our proprietary data?
Yes. We fine-tune Llama, Mistral, and Qwen variants with LoRA or QLoRA on Modal or your own GPUs, distill expensive frontier-model traces into cheaper production models, and run the eval work to prove the fine-tuned model is actually better — not just cheaper.
How do you make sure AI features are safe and accurate?
Every customer-facing system ships with an evaluation harness, prompt-injection defenses, PII redaction, content filters, and abuse rate-limiting. Red-team playbooks are part of pre-launch hardening. We refuse to ship AI features that need more determinism than an LLM can give — and we will tell you so.
Will you work with our existing engineering team?
Yes — about half our AI engagements are embedded squads working alongside an internal team. We bring the AI engineering depth, pair with your product and platform engineers, and leave runbooks, evals, and ADRs so the system is maintainable after we wind down.
Do you offer ongoing AI maintenance after launch?
Yes. Most AI clients keep us on a monthly retainer covering model upgrades (Claude 3 → 4 → 5, GPT-4 → 5 transitions), prompt regression hunting, eval-set expansion as new failure modes surface, and the slow steady work of pushing answer quality up while cost comes down.
— From the journal
Latest articles
- Why Mobile Apps Get Rejected by Apple or Google Sep 4, 2026
- How Often Should a Business Website Be Updated? Sep 2, 2026
- What Is an API and How Can APIs Automate Your Business? Sep 1, 2026
- How to Choose the Right Linux Distribution for a Web Server, Website, Mobile App Backend, or Custom Web Software Aug 31, 2026
- IT Documentation Every Business Should Have Before Something Goes Wrong Aug 28, 2026
- How to Manage Windows and Mac Computers in the Same Company Aug 27, 2026
- How to Build a Private AI Knowledge Base for Your Company Aug 26, 2026
- Computer Problems? How SimplyRem Helps With Software, Hardware, and IT Support Aug 24, 2026
- Can Web-Based Software Slow Down or Freeze Your Computer? What Hardware Do You Actually Need? Aug 21, 2026
- Website Malware: How to Detect It, Clean It Safely, and Prevent It From Coming Back Aug 20, 2026