RAG development services are worth judging on what happens to the questions nobody rehearsed. We have shipped one of these systems and we will show you the code.
The difference between a demonstration and something people rely on sits almost entirely in retrieval quality and in what the system does when the corpus is silent.
What RAG Development Services Should Include
Chunking that follows the document. Splitting on the structure already in the document is unglamorous and it is most of the quality.
Retrieval you can inspect. For any answer, which passages came back and in what order. Without that view a wrong answer cannot be diagnosed, and every system produces some.
Citations that resolve. Every claim points at its source passage and a reader can open it.
Refusal when the corpus is silent. The hard part, and the behavior that decides whether people still trust the system in month three.
An evaluation suite. Real questions with agreed answers, scored on whether the cited passage genuinely supports the claim. Fluency and accuracy come apart precisely where it costs you.
The Baseline Comes First
Before anything gets built, we measure what plain keyword search over the same corpus already achieves on your real questions.
That is the number this has to beat, and it is higher than most buyers expect. Where the margin is thin, improving search is the honest recommendation, and it arrives in week one on our time.
Where it is wide, you hold a business case with a figure attached, which is the reason to hire RAG development services rather than a platform.
Where the Money Goes
Not the model. Recurring cost is dominated by how much context each call carries, which is a design decision rather than a vendor one.
Pull twenty passages when four would serve and the bill multiplies by five, while accuracy often drops because the useful passage is buried among near misses.
We measure that, tune it, and show the arithmetic at ten times your opening volume, so running cost is a decision made at scoping rather than a discovery in month four.
Senior People, Scoping Through Evaluation
The engineer who writes your evaluation set writes the retrieval code, which is why the scoring matches the system.
The code, the prompts, the chunking configuration, the index settings and the evaluation data are yours from the first commit. There is no subscription and no dashboard to log into.
You can re-run the evaluation and get the same score we did, and that reproducibility is what RAG development services should be measured on.
Related Services
Whether it works is AI evaluation, which on this kind of system is not optional.
Where an answer has to trigger an action rather than be read, that is AI agents. Where the job is pulling structured fields out of documents rather than answering questions about them, that is document automation.
Tell us the question your people keep asking and where the answer currently lives.