Retrieval-augmented generation can help a language model answer with information from an organisation's own sources. It does not automatically make those sources accurate, access-controlled, current or suitable for the decision at hand.

A useful RAG system is an information product and an operating capability, not only a model connected to a vector database. Its quality depends on the question, source material, permission model, retrieval pipeline, response controls, evaluation set and the people responsible after launch.

Use this checklist with business, data, security and technical owners. For each item, mark Ready, Needs work or Not applicable, then record an owner and the evidence. This page stores no checklist responses.

What “RAG ready” actually means

Readiness does not mean every source is perfect or every risk has disappeared. It means the team has a bounded use case, understands important failure modes, can enforce the required access boundaries, and has a credible way to measure whether retrieval and answers are useful.

Important limitation

RAG can improve grounding and provide source context, but it does not eliminate unsupported answers. High-impact decisions still need appropriate validation, escalation and human accountability.

A small, restricted pilot can be appropriate when some items need work—provided those gaps are explicit and the pilot cannot create unacceptable consequences.

1. Use case and decision boundary

  • Named usersThe initial user group and its information needs are specific.
  • Task definitionThe system supports a bounded task such as finding policy guidance, comparing technical documents or drafting from approved material.
  • Current baselineThe team knows how the task is completed today and can compare time, quality or success.
  • Value hypothesisThere is a measurable reason to improve the task, not only a desire to “add AI.”
  • Unacceptable outcomesThe team has identified advice, disclosure or actions the system must not produce.
  • Human responsibilityUsers know what they must verify and where accountability remains.
  • Escalation pathThe product can route uncertainty or sensitive questions to an appropriate person or source.

A broad goal such as “chat with all company data” is not a safe first use case. Narrow the users, sources, tasks and allowed outcomes until quality can be evaluated meaningfully.

2. Source quality and ownership

  • Authoritative setThe team can identify which repositories and documents should govern an answer.
  • Content ownersEach important source has someone responsible for accuracy and lifecycle.
  • FreshnessUpdate frequency and acceptable staleness are defined for each source class.
  • Duplicates and conflictsSuperseded versions, copies and contradictory policies can be detected or prioritised.
  • Format coveragePDFs, scans, tables, presentations, web pages and other formats have been sampled for extraction quality.
  • Language and terminologyAcronyms, product names and domain language are represented in evaluation questions.
  • Deletion rulesWithdrawn information can be removed from indexes and caches within an acceptable period.

Retrieval cannot repair a knowledge base that has no authority model. If two policies conflict, the system needs metadata and business rules that help prefer the approved, current source—or it should expose the conflict.

3. Identity, permissions and auditability

  • User identityRequests can be tied to an authenticated identity where the source requires it.
  • Source-level accessThe retrieval layer respects the permissions applied to repositories and documents.
  • Fine-grained boundariesDepartments, clients, projects, regions and confidentiality levels are enforced where needed.
  • Role changesJoiner, mover and leaver events propagate to retrieval access promptly.
  • RevocationRemoving source access also removes retrieval access, including derived indexes and caches.
  • Service accessConnectors and indexing jobs use least-privilege credentials with accountable owners.
  • LogsThe team can investigate who asked, what sources were retrieved and which system version responded without over-collecting sensitive content.

Do not filter only after generation. If a user cannot access a document, its content should not enter that user's model context. Test permission boundaries with adversarial and indirect questions, not only normal queries.

4. Ingestion and knowledge lifecycle

StageReadiness questionFailure to test
ConnectionCan sources be read reliably without excessive privilege?Expired credentials, rate limits, unavailable source
ParsingAre headings, tables, lists and page relationships preserved sufficiently?Scanned page, complex table, broken encoding
SegmentationDo chunks preserve enough context for the target questions?Answer split from condition or exception
MetadataAre owner, date, version, type and permissions attached and filterable?Missing or inconsistent fields
VersioningCan current and superseded content be distinguished?Two policies with similar titles
UpdateHow quickly do additions and changes become searchable?Partial failure or stalled job
DeletionCan removed content disappear from retrieval and relevant caches?Deleted document remains answerable

Record ingestion failures and expose their effect. A pipeline that silently skips documents creates false confidence about what the system “knows.”

5. Retrieval quality

  • Representative questionsThe test set contains real wording, shorthand, ambiguous requests and difficult cases.
  • Relevance labelsDomain reviewers identify which passages should answer each test question.
  • FiltersMetadata, date, jurisdiction, product or permission filters are applied when the task requires them.
  • RankingThe team measures whether the needed evidence appears within the context actually sent to the model.
  • No-answer casesThe set includes questions that the approved sources cannot answer.
  • CitationsUsers can open a stable source location and inspect the supporting context.
  • TerminologyQueries using acronyms, aliases and misspellings are tested.

Evaluate retrieval separately from generation. A fluent response can conceal missing evidence, while poor wording can make good retrieval look bad. Useful measures may include whether relevant evidence was retrieved, its rank, permission correctness and citation validity.

6. Generation boundaries and response design

  • Task instructionThe model is told what it may do, which evidence to use and when to decline.
  • Context treatmentRetrieved content is treated as data, not trusted system instructions.
  • Grounding behaviourAnswers distinguish source-supported facts from interpretation or uncertainty.
  • FallbackThe interface handles insufficient, conflicting or inaccessible evidence clearly.
  • Structured outputWhere another system consumes the response, format and values are validated before use.
  • Sensitive actionsHigh-impact actions require deterministic checks and appropriate approval.
  • User guidanceThe interface explains scope, citations, limitations and how to report a problem.

Prompt instructions are one control, not the entire security model. Enforce important boundaries in identity, retrieval, application logic and downstream validation.

7. Evaluation before and after launch

Create a versioned regression set before tuning on anecdotes. Include common questions, rare but important questions, permission tests, conflicting sources, missing-answer cases and attempts to override instructions.

AreaWhat to measureWho should judge
RetrievalRelevant evidence found, rank, filter and permission correctnessDomain and technical reviewers
GroundingClaims supported by cited context; uncertainty representedDomain reviewers
Task successWhether the answer helps the user complete the defined taskRepresentative users
SafetyRestricted disclosure, instruction attacks, harmful reliance and escalationSecurity, risk and domain owners
ExperienceLatency, source usability, clarity and correction workflowUsers and product owner
EconomicsCost per useful task, usage pattern and budget varianceProduct and finance owners

Set acceptance thresholds appropriate to the use case. A research aid and an automated decision have very different risk profiles. Re-run the set when sources, retrieval logic, models, prompts or permissions change.

8. Privacy and security review

  • Data classificationThe information allowed and prohibited in the system is explicit.
  • Provider termsTraining use, retention, logging, subprocessors and service controls have been reviewed.
  • Data locationRequired processing and storage regions are understood and supported.
  • SecretsCredentials, keys and sensitive configuration are excluded from sources and logs.
  • EncryptionTransport and stored data protections match the organisation's requirements.
  • Threat analysisPrompt injection, data exfiltration, poisoned sources and excessive agency are considered.
  • RetentionQuestions, responses, feedback and traces have defined retention and access rules.
  • Incident responseOwners can contain, investigate and communicate an AI-related incident.
  • Project-specific obligationsLegal, contractual and regulatory requirements have been reviewed by qualified owners.

This checklist is planning guidance, not a compliance determination. Claims about compliance require evidence for the actual design, configuration, providers, data and operating process.

9. Ownership, monitoring and change control

  • Product ownerOne person owns the use case, priorities and acceptance criteria.
  • Knowledge ownersSource teams own accuracy, conflict resolution and retirement.
  • Technical operationSomeone monitors ingestion, retrieval, application health and provider dependencies.
  • Quality signalsDashboards cover failures, latency, feedback and evaluation—not only usage.
  • Feedback triageUser reports become reproducible cases, fixes and regression tests.
  • Change recordSource, model, prompt, retrieval and configuration versions can be traced.
  • Budget controlsUsage limits, alerts and degradation behaviour are defined.
  • RecoveryThe service can be disabled or rolled back without blocking essential work.

Do not rely on a thumbs-up rate alone. Users may reward a confident answer even when its evidence is wrong. Combine behaviour, explicit feedback, sampled review and controlled evaluation.

10. Restricted pilot and expansion criteria

  1. Start with a bounded audience and source set. Choose users who understand the domain and will report issues.
  2. Train users on scope. Explain appropriate questions, citations, verification and escalation.
  3. Observe real tasks. Compare with the baseline and capture new failure modes for evaluation.
  4. Keep high-impact action separate. Do not add automation merely because question answering performs well.
  5. Review at a defined point. Decide to expand, revise, pause or stop using documented evidence.
  6. Expand one risk boundary at a time. New users, sources, actions and regions can each change the threat and quality profile.
Suggested expansion gate

The pilot meets its task-success threshold, permission tests pass, material failures have owners, operating cost is understood, users follow the verification workflow, and security/privacy owners accept the next scope.

From checklist to pilot plan

Summarise the assessment in four lists: ready foundations, gaps that block a pilot, risks that can be constrained within a pilot, and questions requiring evidence. Then define the smallest evaluation that can test retrieval, usefulness and safety with representative users.

RadXsoft designs and builds retrieval systems, AI assistants and workflow automation around real operating constraints. Our team is registered in Varanasi, operates primarily from Delhi and delivers remotely across India and international markets.

Planning a private-data RAG pilot?

Share the target users, task, source systems and access model. We can help turn the readiness gaps into a practical evaluation and delivery plan.