Services › AI in Your Product › Custom RAG systems
AI in Your Product
Custom RAG development: accurate answers from large document sets, with sources and access control
We build retrieval-augmented generation (RAG) systems that answer questions from your documents, manuals, contracts or product data — with a citation for every answer and the same permissions your users already have. Your product calls it as a service.
What is a RAG system?
Retrieval-augmented generation (RAG) is a way of making a language model answer from your own documents: the system first finds the passages relevant to a question, then gives only those passages to the model to write the answer, with links back to the sources. The model doesn't need to be retrained on your data, the documents can change every day, and when the answer isn't in the documents, the system can say so instead of guessing.
Do you need a custom RAG system?
- "Our product has thousands of documents, and users can't find the answer they need."
- "We tried a chatbot over our docs, and it made things up."
- "Different customers must only ever see their own documents."
- "Our documents are technical — tables, part numbers, legal clauses — and generic search misses them."
- "We want AI answers inside our own product, not in a separate tool."
What do we build?
- Ingestion of PDFs, web pages, help centres, databases and file storage, kept in sync as documents change.
- Document processing that keeps structure — headings, tables, page numbers — so answers can point to the exact place.
- Retrieval that combines meaning-based (vector) search with keyword search and re-ranking, tuned on your real questions.
- Access control so each user or customer only retrieves documents they're allowed to see.
- Answers with citations, and a clear "I couldn't find this in the documents" when the answer isn't there.
- An API your product calls, plus an admin view of questions, answers and sources for review.
How does a RAG system work?
- Ingest. Documents are split into passages and indexed for search, and the index updates when documents change.
- Retrieve. A question is matched against the index to find the most relevant passages.
- Filter. Passages the user isn't allowed to see are removed before anything reaches the model.
- Generate. The model writes an answer using only those passages.
- Cite. Each answer links to the passages it came from, so users can check.
Custom RAG or an off-the-shelf tool?
| Off-the-shelf chat-with-docs tool | Custom RAG system | |
|---|---|---|
| Setup | Days | Weeks |
| Where users use it | In the vendor's interface | Inside your product, through an API |
| Per-customer permissions | Limited | Designed around your permission model |
| Complex documents (tables, specs) | Generic handling | Processing tuned to your documents |
| Best when… | One team, internal use, standard documents | Customer-facing, large or specialised, multi-tenant |
For an internal assistant over company documents, our knowledge assistant service is usually the faster route.
How do you make RAG answers accurate?
Most RAG failures are retrieval failures: the right passage was never found, so the model answered from the wrong one. That's why we spend most of the effort on document processing and retrieval, not on the prompt.
We build a test set of real questions with known answers from your documents, measure how often the right passage is retrieved and how often the answer is correct and cited, and re-run that test whenever documents, models or settings change. The system is told to answer only from the retrieved passages and to say when it can't.
What do we build it with?
We choose components to fit your data, scale and hosting requirements. For example:
Where data must stay on your infrastructure, the whole system can run with privately hosted models.
What about data privacy and security?
Permissions are enforced at retrieval, before any text reaches the model, so a clever question can't pull out another customer's documents. Documents and the search index can stay in your own database or cloud account, keys stay on the server, and the model provider — or a privately hosted model — is agreed with you in the proposal.
How long does it take and what does it cost?
A first custom RAG system typically takes 4–8 weeks, depending on the volume and variety of documents and the permission model. Setup starts at $3,000, and monthly care starts at $300 per month. Model usage and hosting are paid to the providers and grow with documents and questions; we estimate them with you up front.
Monthly care covers monitoring answer quality, re-running the test set, adding new document sources and adjusting retrieval. Payments are milestone-based, and the launch includes one month of free bug fixing.
When is a custom RAG system not a fit?
Good fit when…
- Answers must come from a large or specialised document set
- Users need to see where each answer came from
- Different users may only see different documents
- AI answers belong inside your own product
Not a fit when…
- You have a few dozen documents and one team using them
- Answers need calculations across a database, not text
- The documents are out of date or contradict each other
- A standard chat-with-docs tool already does the job
Frequently asked questions
What is retrieval-augmented generation (RAG)?
RAG is a method where a system first searches your documents for passages relevant to a question, then gives only those passages to a language model to write the answer, with citations. It lets a model answer from your own, up-to-date data without retraining it.
Is RAG better than fine-tuning a model?
For answering questions from documents, usually yes. RAG works with documents that change, shows its sources and respects permissions. Fine-tuning is better for teaching a model a style or format, and the two can be combined.
How do you stop a RAG system from making things up?
By improving retrieval so the right passage is found, telling the model to answer only from retrieved passages, showing citations, and letting the system say when an answer isn't in the documents. Accuracy is measured on a test set of real questions before launch.
Can each customer only see their own documents?
Yes. Permissions are applied when passages are retrieved, before anything reaches the model, so a user can only get answers from documents they are allowed to see.
How much does a custom RAG system cost?
At Webziper, a custom RAG system starts at $3,000 to set up, with monthly care from $300 per month. Model usage and hosting depend on the number of documents and questions.
Users can't find answers in your documents?
Send us a sample of your documents and the questions people ask. We'll tell you how well RAG would work on them.
Book a Free Call →