Services › AI in Your Product › Custom RAG systems

AI in Your Product

Custom RAG development: accurate answers from large document sets, with sources and access control

We build retrieval-augmented generation (RAG) systems that answer questions from your documents, manuals, contracts or product data — with a citation for every answer and the same permissions your users already have. Your product calls it as a service.

What is a RAG system?

Retrieval-augmented generation (RAG) is a way of making a language model answer from your own documents: the system first finds the passages relevant to a question, then gives only those passages to the model to write the answer, with links back to the sources. The model doesn't need to be retrained on your data, the documents can change every day, and when the answer isn't in the documents, the system can say so instead of guessing.

Do you need a custom RAG system?

What do we build?

How does a RAG system work?

1Ingest documents split and indexed2Retrieve relevant passages found3Filter only what the user may see4Generate answer from those passages5Cite sources linked1Ingest documents split and indexed2Retrieve relevant passages found3Filter only what the user may see4Generate answer from those passages5Cite sources linked
  1. Ingest. Documents are split into passages and indexed for search, and the index updates when documents change.
  2. Retrieve. A question is matched against the index to find the most relevant passages.
  3. Filter. Passages the user isn't allowed to see are removed before anything reaches the model.
  4. Generate. The model writes an answer using only those passages.
  5. Cite. Each answer links to the passages it came from, so users can check.

Custom RAG or an off-the-shelf tool?

Off-the-shelf chat-with-docs toolCustom RAG system
SetupDaysWeeks
Where users use itIn the vendor's interfaceInside your product, through an API
Per-customer permissionsLimitedDesigned around your permission model
Complex documents (tables, specs)Generic handlingProcessing tuned to your documents
Best when…One team, internal use, standard documentsCustomer-facing, large or specialised, multi-tenant

For an internal assistant over company documents, our knowledge assistant service is usually the faster route.

How do you make RAG answers accurate?

Most RAG failures are retrieval failures: the right passage was never found, so the model answered from the wrong one. That's why we spend most of the effort on document processing and retrieval, not on the prompt.

We build a test set of real questions with known answers from your documents, measure how often the right passage is retrieved and how often the answer is correct and cited, and re-run that test whenever documents, models or settings change. The system is told to answer only from the retrieved passages and to say when it can't.

What to expect: a RAG system can be very reliable on questions your documents answer clearly. It can't fix documents that contradict each other or are out of date — it will surface those problems instead.

What do we build it with?

We choose components to fit your data, scale and hosting requirements. For example:

PostgreSQL + pgvectorQdrantPineconeOpenAIAnthropic ClaudeGoogle GeminiOpen-source modelsn8nBubble.io and custom backends

Where data must stay on your infrastructure, the whole system can run with privately hosted models.

What about data privacy and security?

Permissions are enforced at retrieval, before any text reaches the model, so a clever question can't pull out another customer's documents. Documents and the search index can stay in your own database or cloud account, keys stay on the server, and the model provider — or a privately hosted model — is agreed with you in the proposal.

How long does it take and what does it cost?

A first custom RAG system typically takes 4–8 weeks, depending on the volume and variety of documents and the permission model. Setup starts at $3,000, and monthly care starts at $300 per month. Model usage and hosting are paid to the providers and grow with documents and questions; we estimate them with you up front.

Monthly care covers monitoring answer quality, re-running the test set, adding new document sources and adjusting retrieval. Payments are milestone-based, and the launch includes one month of free bug fixing.

When is a custom RAG system not a fit?

Good fit when…

  • Answers must come from a large or specialised document set
  • Users need to see where each answer came from
  • Different users may only see different documents
  • AI answers belong inside your own product

Not a fit when…

  • You have a few dozen documents and one team using them
  • Answers need calculations across a database, not text
  • The documents are out of date or contradict each other
  • A standard chat-with-docs tool already does the job

Frequently asked questions

What is retrieval-augmented generation (RAG)?

RAG is a method where a system first searches your documents for passages relevant to a question, then gives only those passages to a language model to write the answer, with citations. It lets a model answer from your own, up-to-date data without retraining it.

Is RAG better than fine-tuning a model?

For answering questions from documents, usually yes. RAG works with documents that change, shows its sources and respects permissions. Fine-tuning is better for teaching a model a style or format, and the two can be combined.

How do you stop a RAG system from making things up?

By improving retrieval so the right passage is found, telling the model to answer only from retrieved passages, showing citations, and letting the system say when an answer isn't in the documents. Accuracy is measured on a test set of real questions before launch.

Can each customer only see their own documents?

Yes. Permissions are applied when passages are retrieved, before anything reaches the model, so a user can only get answers from documents they are allowed to see.

How much does a custom RAG system cost?

At Webziper, a custom RAG system starts at $3,000 to set up, with monthly care from $300 per month. Model usage and hosting depend on the number of documents and questions.

Users can't find answers in your documents?

Send us a sample of your documents and the questions people ask. We'll tell you how well RAG would work on them.

Book a Free Call →