Preparing data for AI starts with deciding which information is trustworthy, current and appropriate for the task. A large upload is not a knowledge strategy. The aim is a manageable collection with clear ownership, permissions and a way to correct mistakes.
Inventory sources by usefulness
For each source, record its owner, audience, update frequency, format and access rules. Identify the questions it should answer. A small collection of current procedures can be more useful than a shared drive containing years of contradictory drafts.
Ask users for the documents they actually consult. Include difficult formats such as scanned forms and tables in the assessment. Test extraction quality early; if a critical column is lost, later model improvements may not repair the underlying evidence.
Resolve versions and contradictions
Mark the authoritative version and define what happens to superseded material. Preserve dates and identifiers that help users verify an answer. If two departments have different policies, record the audience rather than merging them into one apparently universal rule.
Use AI to help flag duplicates or inconsistent wording, but have the source owner approve corrections. A model should not decide which policy is valid because one document sounds more confident or has a newer file timestamp.
Preserve access boundaries
Classify confidential and personal information and minimise what is necessary for the pilot. Map permissions from the source to the retrieval or application layer. Test a user with limited access and a user whose access has been revoked.
Derived copies matter too: extracted text, indexes, embeddings, logs and backups can contain sensitive information. Include them in retention and deletion processes. If a source cannot provide reliable permission metadata, restrict its use until a safe access model is defined.
Make updates an operating process
Choose how changes enter the system and how quickly they should appear. Record failed imports and give owners a way to see whether their latest document is available. A silently stale index can produce plausible answers long after the source has changed.
Test deletion and correction. Remove a document, rerun a relevant question and confirm that it is no longer used. Check that old excerpts do not survive through another copy or cache.
Prepare an evaluation collection
Pair real questions with expected evidence and acceptable answers. Include missing information and ambiguous language. Where Arabic and English are used, test both and check whether translated terminology changes meaning.
Keep this collection with the source owners and the delivery team. It turns data preparation into a visible quality process and gives you a basis for evaluating model, retrieval and document changes over time.
Sources & further reading
Primary references for the technical background and regional statements in this guide. Planning examples and checklists are Novacom’s practical guidance; examples are illustrative unless explicitly identified as project experience.