Jabrod Services

We build production RAG and LLM agent systems — and cut what they cost you to run.

Done-for-you engineering from the team building Jabrod. Fixed scope, fixed price, no discovery phase and no retainer minimum. If you'd rather build it yourself, the product is right here.

80,000+

users served by AI systems we've built and run in production

40%

cut from LLM inference spend on a live platform, without quality loss

2 wks

from kickoff to a working, evaluated retrieval pipeline

Services

Three things, priced up front.

Most AI work is sold as an open-ended engagement with a discovery phase attached. These are fixed. You know the scope, the price and the date before you commit to anything.

Start here

LLM Cost Audit

$500fixed

1 week · remote · no production access required

Most teams running LLMs in production are spending 20–40% more than they need to. It's rarely one big thing — it's uncached prompts, a frontier model doing work a small one could do, and retrieval stuffing far more context into every call than the answer actually needs.

  • A breakdown of where your spend actually goes, by route and by call type
  • The specific waste, quantified — caching, model selection, context size, retries
  • A ranked list of fixes with estimated savings and effort for each
  • A 45-minute walkthrough call, and the document is yours to keep either way
Under 20% found → you don't pay
Book the audit

RAG Pipeline Sprint

$2,000fixed

2 weeks · fixed scope

A working retrieval system over your documents, answering questions with citations you can trace back to the source — parsing, chunking, embeddings, retrieval and evaluation, deployed to your infrastructure.

  • The full pipeline, running, in your stack
  • An evaluation set so you can prove it works — and catch it when it stops working
  • Deployment and maintenance documentation written for your team, not for us
Talk it through

RAG Eval & Repair

$1,200fixed

10 days · fixed scope

You already have a retrieval system. It gives confident, wrong answers and nobody can say exactly when or why. This finds out.

  • An eval harness with a labelled set drawn from your real queries
  • A diagnosis of where retrieval breaks — chunking, embeddings, ranking or the prompt
  • The fixes shipped, with before-and-after numbers
Talk it through

These are done-for-you engagements, not product plans. Looking for Jabrod's own pricing? That's over here. Prices in USD; UK clients are quoted and invoiced in GBP. Larger builds and monthly retainers are quoted after a call.

Selected work

Numbers, not adjectives.

40%

Cutting inference cost on a live legal-AI platform

The problem
A production legal assistant serving 80,000+ users was routing every query, trivial lookups included, through a frontier model, with no caching layer and retrieval returning far more context than answers required.
What we did
Introduced prompt caching across the highest-volume paths, and built a routing layer that classifies each query and sends the simple majority to lighter models, keeping the frontier model for work that genuinely needs it.
Result
A 40% reduction in inference spend, with no measurable drop in answer quality.

4-stage

A cited retrieval pipeline at production scale

The problem
Legal answers are worthless if users can't check them. The system needed to ground every response in source documents and survive real traffic.
What we did
Designed a four-stage pipeline (document parsing, chunking, embedding, retrieval) behind a multi-agent routing layer with explicit reasoning stages that dispatch each query to the right tools and workflows.
Result
Context-grounded, citable answers running in a single production system alongside auth, background jobs and frontend delivery.

Who you'll work with

Small on purpose.

The person who scopes your project is the person who writes the code and the person who answers your emails. There's no account manager and nothing gets handed off.

Anshuman Tiwari

Founder, Jabrod

I build production AI systems — retrieval pipelines, agent workflows, and the full stack around them. Day to day I work on a Next.js AI platform serving 80,000+ users, where I designed the retrieval pipeline and the multi-agent routing layer, and took 40% out of the inference bill.

Further down the stack, I've implemented FlashAttention-2 forward and backward attention kernels in Triton from the original papers — online softmax with running statistics, O(N) activation memory, benchmarked against PyTorch's own implementation. It's not what most clients need, but it's why I can tell you what your model is actually doing when the bill arrives.

Languages
Python · TypeScript · C++
AI
RAG · LLM agents · LangChain · LangGraph · evals · prompt caching · vector databases · PyTorch · Triton
Platform
Next.js · React · Node · FastAPI · PostgreSQL · MongoDB · Docker · AWS · GCP

Start with the audit.

Twenty minutes, no obligation, and you'll leave the call knowing roughly what your LLM spend should be — whether or not you hire us.