Semitora.

30 June 2026 · Updated: 10 September 2026

Data readiness for RAG — a checklist before you deploy AI

Many enterprise AI projects don’t fail on the model — they fail on the data. Data readiness for RAG is the state where your documents are complete, clean, governed by permissions and kept up to date — fit to be the knowledge base your AI answers from, with a source. The model is rarely the edge; what matters is what you feed it. Walk this checklist — 25 questions across six areas — before you build, not when the prototype is already hallucinating.

If you’re still working out what RAG is, start with RAG on company documents. Below we assume you know what it’s for, and you’re asking the next question: is my data ready for it.

1. Sources and scope

Before you index anything, you need to know what you’re indexing and where it lives.

Before evaluation, compare the agreed set of expected documents with processed, rejected and partial documents; confirm that the processed content is in the index. The reconciliation report is for an authorized collection owner or administrator. Below is a hypothetical ingestion manifest, with no customer data or results.

Field Example
Document / version DOC-C / v2
Ingestion result Partial
Omitted scope Pages 4–5
Reason Incomplete OCR
Owner Collection owner B

A deliberate scope exclusion is not a processing error such as incomplete OCR; record it separately, outside the set expected for ingestion. The permissions filter (ACL) is a valid access boundary, not an ingestion error. Do not reveal document names or identifiers outside the user’s permissions, including in this kind of manifest.

Ingestion and index coverage are not the same as retrieval top-k or citations for a specific answer. The model does not see the entire knowledge base: it gets selected passages accessible to that user. Missing important sources can distort evaluation results; they do not necessarily improve them. Before accepting evaluation results, the collection owner must explicitly decide whether to fill the gaps and repeat the tests or accept a limited scope with the impact of the gaps documented. The go/no-go verdict stays in the Evidence Pack.

2. Permissions and sensitive data

This is the area that most often derails a project after the fact — and is the hardest to fix once you’re live.

3. Document quality and structure

The model is only as good as the chunk it gets. Garbage in, garbage in the citation.

4. Freshness and versioning

A knowledge base isn’t a one-day snapshot. Data that doesn’t refresh ages faster than you think.

5. Tests and quality metrics

Without tests, “it works” is a hunch, not a fact. Building isn’t enough — you have to measure.

6. Cost and maintenance

The most expensive part of GenAI is usually not inference but data engineering — and it doesn’t end at go-live.

In short

Data readiness for RAG is checked across six areas: sources and scope, permissions and sensitive data, document quality and structure, freshness and versioning, quality tests, and cost and maintenance. If you answer “I don’t know” to most of the questions, that isn’t a reason to drop AI — it’s the first phase of the project. The cheapest time to find out is before you build, not after.

What next

How we build RAG on company documents — with sources, on AWS — is on the RAG / knowledge bases page. Tidying data (ETL) and the knowledge base are a distinct delivery step for us, described in how we work. If you don’t know where to start, start with an audit: we’ll walk this checklist on your data.