On-premises RAG keeps the selected retrieval and generation components inside infrastructure your organisation operates. It can suit organisations whose documents cannot be sent to an external model service. The important question is whether the entire data path stays inside the approved boundary, including logs, embeddings and support access.
- 01Approved sources
- 02Ingestion & access metadata
- 03Search index
- 04Retrieval API
- 05Private model endpoint
Illustrative process. The exact components and controls depend on the workload.
Draw the data path first
A reference architecture has five stages: approved sources, an ingestion worker, a searchable index, a retrieval API and a model-serving endpoint. The user enters through an application connected to company identity. The retrieval layer applies permissions before returning evidence to the model.
An internal document does not become harmless when converted to an embedding. Treat indexes, extracted text, temporary files and conversation records according to their content. Identify which components can reach the internet and which packages or model weights need an approved import process.
Build permissions into retrieval
Consider a procurement assistant with general purchasing guidance and restricted supplier negotiations. An employee allowed to read the guidance must not see passages from the negotiation folder. Filtering only the final response is too late: the restricted information should never enter that user’s model context.
Record source identifiers and access metadata during ingestion. Define how permission changes and deletions propagate. Test revoked access, nested groups and shared links, not just two static test accounts. Where a connector cannot preserve permissions reliably, limit the pilot to a deliberately shared collection.
Size the whole system
Model inference is only one workload. Document extraction, OCR, embedding generation, indexing and backups also consume capacity. Separate initial ingestion throughput from everyday query latency. A large document import should not make employee conversations unusable.
Benchmark representative questions with realistic passage sizes and concurrent users. Include Arabic and English material if both are required. Tables, forms and mixed-language documents need their own tests; a good answer on plain English paragraphs proves little about those cases.
What to prepare before a pilot
- A named document owner and a small, approved source collection.
- Identity groups, retention rules and a deletion test.
- Expected answers and citations for real staff questions.
- Network, model-licence, hardware and backup requirements.
- An operating owner for upgrades, incidents and evaluation after changes.
A useful acceptance test includes a restore from backup and an unavailable model endpoint. Users should receive a clear failure message and a route to the original document, rather than a fabricated answer.
Is on-premises automatically cheaper?
No. Compare the full operating cost: hardware utilisation, staff time, power, resilience and upgrade effort. Private deployment is often chosen for control or integration requirements. Validate the financial case separately rather than assuming that owning GPUs makes each useful answer less expensive.
Sources & further reading
Primary references for the technical background and regional statements in this guide. Planning examples and checklists are Novacom’s practical guidance; examples are illustrative unless explicitly identified as project experience.