RAG supplies information when a model answers. Fine-tuning changes a model using training examples. Start with the failure you need to fix: missing company knowledge, inconsistent behaviour or a task the model cannot perform reliably. Those are different engineering problems.
| Need | Start by evaluating | Why |
|---|---|---|
| Current company facts | RAG / authorised lookup | Supply evidence when answering. |
| Consistent task behaviour | Prompting, then fine-tuning if justified | Measure a repeatable behaviour gap. |
| Exact calculations | Validated code or tools | Check the inputs and result. |
| Permissions | Application access controls | Neither retrieval nor training replaces authorisation. |
Choose by the failure, not the fashion
If your assistant cannot find yesterday’s approved policy, improve retrieval and source freshness. Training the policy into model weights creates an awkward update process and does not provide document-level access control.
If a model repeatedly misclassifies your organisation’s specialist support categories despite clear instructions and examples, fine-tuning may be worth evaluating. You will need representative labelled cases, an independent test set and a way to measure whether the adapted model actually improves the workflow.
If the underlying source is wrong, neither approach repairs it. If the question requires an exact calculation, use validated code or a suitable tool. Do not treat every failure as evidence that you need a larger model.
A procurement example
Suppose staff need to draft a purchase request in a consistent format and explain the current approval policy. Retrieval can supply the applicable policy. A prompt and a validated output schema may handle the format. Only consider training if a persistent, measurable behaviour gap remains.
The system can use both techniques: a model adapted to a specialist task can still retrieve current evidence. Evaluate the combination against a simpler baseline so the added infrastructure and training maintenance have a clear reason to exist.
Compare the ongoing work
RAG requires source ownership, ingestion, retrieval tuning and permissions. Fine-tuning requires dataset preparation, training, evaluation and management of model versions. Both require monitoring and regression checks after changes.
For a practical comparison, run the same held-out tasks through your current process, a prompt-only baseline, retrieval and any proposed adapted model. Count accepted outputs, critical mistakes and human correction time. Measure total cost per accepted task, including retries and review, rather than comparing a single API price.
Questions to answer before choosing
- Does the information change frequently, and must the answer cite a source?
- Do different users have different rights to the underlying information?
- Is the desired behaviour clear enough for people to label consistently?
- Do you have examples that represent real edge cases, not just easy successes?
- Who maintains the documents or training data after launch?
For many internal knowledge assistants, begin with a retrieval pilot and a strong evaluation set. That recommendation is about the job being done; it is not a claim that fine-tuning has become obsolete.
Sources & further reading
Primary references for the technical background and regional statements in this guide. Planning examples and checklists are Novacom’s practical guidance; examples are illustrative unless explicitly identified as project experience.