RAG development services for answers you can check
We build retrieval-augmented generation (RAG) systems that answer from your documents and live records, cite where each answer came from, and are measured on a test set before anyone relies on them.
MappOptimist builds retrieval-augmented generation (RAG) systems for enterprises and SaaS teams: ingestion and chunking, embeddings, hybrid search over vector stores, real-time retrieval from live systems, grounded answers with citations, and evaluation of retrieval and answer quality. RAG can power a standalone assistant or the knowledge layer of an AI agent, delivered as a project or by LLM engineers who join your team.
Problems we solve
The model answers confidently and wrongly
A general-purpose LLM does not know your policies, products or customers. Without retrieval it fills the gaps with plausible text, and nobody can tell which answers to trust.
Search finds the document but not the answer
Keyword search returns the right PDF and leaves people to read forty pages. Plain vector search finds similar wording but misses exact terms such as product codes, clause numbers and names.
Knowledge is spread across systems
Answers depend on documents in one place, tickets in another and live records in a database. A pilot that indexed one folder does not cover the questions people actually ask.
No one can say if it is getting better or worse
Prompts, chunk sizes and models change, and quality is judged by trying a few questions. Without a test set, a change that helps one answer quietly breaks ten others.
What we build
Ingestion and chunking pipelines
Connectors for your document stores, wikis, tickets and databases, with parsing, cleaning and chunking tuned to how each source is written, plus metadata such as owner, date and access level on every chunk.
Hybrid search and re-ranking
Vector search combined with keyword search, so meaning and exact terms both count, then re-ranking to put the most useful passages first. Vector stores such as Pinecone, Weaviate or Qdrant, chosen for your hosting and scale.
Real-time retrieval from live systems
For questions about orders, accounts or status, the system queries live APIs and databases rather than a stale index. Our multi-agent support bot for a logistics and e-commerce platform answers tracking and payment queries this way.
Grounded answers with citations
Answers are generated only from retrieved passages and link back to them, so a reviewer can check the source in one click. When retrieval finds nothing relevant, the system says so instead of guessing.
Access control on retrieval
Users only retrieve what they are allowed to see. Permissions are applied when passages are fetched, not after an answer is written, so restricted content never reaches the model for the wrong person.
Evaluation and monitoring
A test set built from real questions scores retrieval and answers before release and after every change. In production we trace each query and track quality, latency and cost.
How an engagement runs
- Step 01
Collect real questions
We gather the questions people actually ask, with the right answers and where they live. This becomes the evaluation set and decides which sources are indexed first.
- Step 02
Build the retrieval layer
Ingestion, chunking, embeddings, hybrid search and access control for the first sources, measured against the evaluation set until retrieval finds the right passages.
- Step 03
Add generation and citations
Grounded answers with citations and a clear "not found" path, then a pilot with real users whose feedback flows back into the test set.
- Step 04
Launch and tune
Deploy into your cloud, monitor quality, latency and cost, and add sources in batches, each one checked against the evaluation set before it goes live.
Ways to work with us
RAG build
We scope and deliver a working RAG system end to end, from source connectors to a deployed assistant or API, against milestones and quality targets agreed up front.
Knowledge layer for an AI agent
Retrieval built as one part of a larger agent system, alongside tool calling, approvals and orchestration. See our AI agent development service.
LLM engineers for your team
LLM engineers who join your team on a monthly rolling basis to build or improve your retrieval pipeline. You interview and approve everyone before they start.
Need individual engineers rather than a project? See the roles you can hire and their rates.
Technology we work with
Where our retrieval work runs
Retrieval in these systems feeds AI agents: live data for a support bot, and a vector store behind an enterprise agent platform.
Autonomous support agents
A multi-agent customer support chatbot — specialised bots for tracking, support and payments — that automates repetitive queries and frees the team for complex issues.
Read the case study →AI Agents · EnterpriseAgentic AI Platform
Goal-driven AI agents that plan, call tools and complete real workflows under human oversight.
Read the case study →Frequently asked questions
What is RAG, and when do we need it?
Retrieval-augmented generation fetches relevant passages from your own documents or systems and gives them to the model, which answers from them and cites them. You need it when answers depend on information the model was not trained on, such as your policies, products or customer records, or when people must be able to check where an answer came from.
RAG or fine-tuning?
Use RAG when the knowledge changes or must be cited: new documents are indexed in minutes and each answer points to its source. Fine-tuning changes how a model writes or behaves, not what it knows, and is retrained when the facts change. Most business assistants need RAG first; fine-tuning is occasionally added for tone or format.
How do you measure whether a RAG system works?
We build a test set of real questions with known answers and sources. Retrieval is scored on whether the right passages come back, and answers on whether they are correct, grounded in those passages and properly cited. Every change to chunking, models or prompts runs against the set before release.
Can RAG respect who is allowed to see what?
Yes. Access levels are stored with each chunk and applied at retrieval time, so a user's query only searches content they are permitted to read. This matters for HR, legal and customer data, where filtering after the answer is written would already be too late.
Which vector database should we use?
It depends on where the data must live, how much of it there is and what your team already runs. We work with Pinecone, Weaviate and Qdrant, and we combine vector search with keyword search, because exact terms such as codes and names are often what the question hinges on.
What does RAG development cost?
Cost depends mainly on how many sources must be connected, how messy the documents are, access-control needs and evaluation depth. Projects are quoted after a short scoping exercise. For team augmentation, our 2026 rate card lists an LLM Engineer at $40/hr (2-3 years), $48/hr (4-6 years) and $55/hr (6+ years). Rates are indicative and confirmed per engagement.
Ground your AI in your own data
Tell us which questions people need answered and where the answers live today. We reply within one business day with how we would build the retrieval layer.