Relevant evidence is not retrieved
Useful passages may be missed because of chunking, metadata, filters, query rewriting, indexing, ranking or reranking choices.
Get an independent, evidence-led view of how your retrieval-augmented generation system finds knowledge, uses evidence, handles uncertainty and operates within quality, security and governance boundaries. The assessment turns observed gaps into a prioritised remediation and retest plan.
Assessment outcomes apply to the agreed evidence, scenarios and environment. They do not guarantee error-free answers, certification, regulatory compliance or elimination of residual AI risk.
A quality assessment helps distinguish whether the material issue sits in the source corpus, document processing, retrieval, context assembly, generation, access controls or operating process before teams make expensive changes in the wrong layer.
Useful passages may be missed because of chunking, metadata, filters, query rewriting, indexing, ranking or reranking choices.
A response can sound credible while extending beyond retrieved evidence, omitting material conditions or citing passages that do not support the claim.
Source links may be present yet point to stale, incomplete, conflicting or insufficiently authoritative evidence.
Identity-aware filtering can break across indexes, caches, tools, shared stores or multi-tenant contexts and expose information outside intended boundaries.
Model, prompt, corpus, embedding, chunking or reranking changes can improve one scenario while quietly degrading another.
Teams may have dashboards and demos without agreed acceptance criteria, risk owners, evidence retention, retest rules or residual-risk decisions.
Share the business use case, current RAG architecture and the quality concern you need to resolve. We can help define a bounded assessment around the evidence and decision that matter.
The engagement is a structured independent review under the AI Assessments family. It is designed to expose material gaps, improve decision quality and create a practical remediation path without turning a limited evidence set into an unsupported certification claim.
DataConsultant first clarifies the use case, users, answer boundaries, risk, environment and decision to be supported. Assessment criteria are then selected for the system under review. Evidence can combine documentation, configuration, traces, representative scenarios, existing metrics, interviews and controlled tests where authorised.
Not every assessment needs every domain. The review concentrates effort where the intended use, observed failures, evidence and risk justify deeper examination.
Authority, freshness, duplication, completeness, structure, permissions, metadata, provenance and lifecycle of the knowledge being retrieved.
Knowledge foundationDocument processing, semantic boundaries, overlap, hierarchy, embeddings, index coverage and whether source meaning remains retrievable.
RetrievabilityQuery transformations, filters, hybrid methods, top-k selection, ranking, reranking, recall, precision and evidence authority.
Search qualityWhether generated claims remain supported by supplied context and whether material expected information is missing from the answer.
Answer qualityWhether citations support the associated claims, expose useful evidence, preserve source identity and help users verify important answers.
Evidence traceUnsupported questions, contradictory sources, out-of-scope prompts, prompt injection, poisoned content, ambiguity and adverse scenarios.
Boundary behaviourPermission-aware retrieval, multi-user boundaries, sensitive data exposure, logging, caching, data movement and control evidence.
Control integrityMonitoring, latency and cost visibility, version traceability, regression controls, issue ownership, escalation, retesting and approval decisions.
Operational assuranceThe assessment connects the user task to system evidence so teams can avoid arguing about a single aggregate score and instead understand where quality degrades and what should change next.
A RAG quality metric is useful only when its inputs, ground truth, evaluator method and decision meaning are understood. The assessment may combine quantitative measures with structured human review and qualitative control evidence.
| Quality question | Possible evidence or measure | What it can help reveal | Important limitation |
|---|---|---|---|
| Did retrieval find useful evidence? | Retrieval relevance judgements, precision@k, recall@k, ranking measures, coverage | Missing passages, weak ranking, filter errors, source gaps or query-treatment issues. | Reliable relevance labels or reference evidence may be required; retrieval quality alone does not prove answer quality. |
| Did the answer stay within the evidence? | Grounding claim-to-context checks, groundedness or faithfulness review | Unsupported claims, synthesis beyond sources and weak abstention when evidence is insufficient. | Automated judges can disagree or miss subtle misinterpretation; human review may be required for material cases. |
| Did the answer cover what mattered? | Completeness reference answers, rubrics, required-point coverage | Omitted conditions, partial answers and responses that are correct but operationally incomplete. | Completeness depends on the quality of reference expectations and can vary by user role or task. |
| Can users verify material claims? | Citations citation precision, coverage, source authority and traceability review | Citations that are present but irrelevant, insufficient, stale or misaligned with the claim. | Citation presence does not prove that a source is authoritative or current. |
| Does the system respect boundaries? | Controls refusal cases, permission tests, injection scenarios, sensitive-data checks | Unsafe answer paths, cross-user leakage, poisoned content sensitivity or weak policy enforcement. | A bounded quality assessment is not a substitute for exhaustive security testing or formal assurance. |
| Can quality be operated reliably? | Operations trace coverage, change records, regression evidence, latency and cost signals | Monitoring blind spots, undocumented changes, missing retest gates and operational trade-offs. | Targets depend on business use, service architecture and contractual or operational requirements. |
Deliverables are scoped around the decision and evidence available. The objective is to leave business, AI, engineering, security and governance teams with a common view of the material issues and a practical route to validation.
Intended use, stakeholders, boundaries, quality questions, evidence needs, assumptions and decision criteria.
Documents, traces, test artefacts, interviews and environments reviewed, plus missing evidence and scope limitations.
Source-to-answer flow showing where quality signals, controls, dependencies and failure modes sit.
Evidence-backed observations across the agreed quality, risk, control and operational domains.
Material gaps with consequence, severity rationale, affected scenarios, evidence and accountable follow-up.
Where supportable, links between failures and source, chunking, metadata, retrieval, prompting, policy or operational conditions.
Recommended actions organised by business impact, risk, dependency, evidence strength, remediation effort and retest need.
Concise decision view, unresolved risks, next actions, owners and what evidence should be rechecked after remediation.
Request an assessment that connects retrieval and answer evidence to business impact, risk, remediation ownership and the next release or investment decision.
Access is agreed during mobilisation and should be proportionate to the assessment. Sensitive material can be minimised, redacted or reviewed in client-approved environments when appropriate.
The sequence can be adapted to the system and evidence available. Timing is confirmed after scoping rather than inferred from a generic package.
Clarify intended use, users, risk, business impact, scope boundaries, stakeholders and what decision the assessment must support.
Output: assessment brief and criteriaUnderstand sources, ingestion, indexes, retrieval, reranking, prompts, models, citations, permissions, interfaces and operating dependencies.
Output: system map and evidence requestInspect documentation, configuration, traces, existing tests, metrics, incidents, policies and stakeholder evidence within authorised scope.
Output: evidence register and limitationsWhere agreed, examine representative retrieval and answer behaviour across normal, difficult, boundary and risk scenarios.
Output: test observations and examplesSeparate symptoms from likely source, retrieval, generation, control or operational contributors and document evidence strength.
Output: findings and risk/gap registerCompare impact, risk, reproducibility, dependency, control strength, remediation effort and retest requirements.
Output: prioritised remediation backlogAlign decision owners on material findings, residual risk, action ownership and evidence required before the next release or review.
Output: executive readout and validation planIf your team already has logs, dashboards or test results but lacks a clear interpretation, the assessment can concentrate on material gaps, root causes, ownership and what should be retested next.
The assessment can use qualitative or quantitative scoring only when the method is justified. The more important objective is to make the severity rationale, affected scenario and validation path clear enough for accountable owners to act.
No single factor decides severity. The assessment combines business impact with technical evidence, risk and practical remediation context.
Release and risk acceptance remain client-owned decisions. The assessment can organise evidence into practical decision bands without claiming a universal certification threshold.
Criteria are selected for the client’s architecture and risk. Current vendor and risk-management guidance can be used as reference material where relevant, while the final assessment remains grounded in the actual business use case and evidence.
Current Microsoft guidance distinguishes retrieval-process evaluation from response-level evaluation and includes concepts such as document retrieval, groundedness, relevance and response completeness.
Review Microsoft guidance ↗AWS documents retrieve-only and retrieve-and-generate evaluation approaches with built-in measures for retrieval context and generated-response quality.
Review AWS guidance ↗NIST AI 600-1 is a voluntary cross-sector companion to the AI Risk Management Framework for managing generative-AI risks across design, development, use and evaluation.
Review NIST publication ↗OWASP documents prompt-injection risks and vector or embedding weaknesses that can affect RAG systems, including access leakage, poisoned content and retrieval manipulation.
Review OWASP guidance ↗Reference frameworks and platform features change over time. Their applicability depends on the selected technology, jurisdiction, sector, contractual obligations and internal policy. The assessment does not replace authorised legal, privacy, security or regulatory interpretation.
No fixed public fee is stated for this service because a focused evidence review and a multi-environment assessment with human evaluation, control testing and retesting have materially different scope. A written quote is prepared after the assessment boundary and evidence needs are understood.
Pricing is based on the agreed assessment question, systems, evidence, testing depth, stakeholders and deliverables. Third-party platform, cloud, model, storage or evaluation-tool consumption is separate from consulting fees unless explicitly included in the proposal.
No market-derived competitor price has been presented as a DataConsultant fee. A numeric figure should only be published when an approved DataConsultant price or sufficiently comparable and supportable market evidence is available.
If the problem is broader implementation, source remediation, formal security testing or continuous evaluation, an adjacent service may be a better starting point.
Tell us what decision is blocked, what evidence you already have and which parts of the RAG stack are in question. We can recommend a focused assessment boundary or a more suitable adjacent service.
RAG quality is rarely owned by one component. An assessment is more actionable when the review can distinguish source-data, retrieval, model, control and operating issues and route remediation to the right accountable team.
Start from the business decision and evidence rather than defending a predetermined platform, model or supplier choice.
Connect retrieval outcomes to source quality, metadata, provenance, permissions, lifecycle and data-management conditions.
Consider quality alongside privacy, security, human oversight, monitoring, change control and accountable decision rights.
Translate findings into owners, dependencies, remediation priorities, retest evidence and adjacent implementation needs.
These answers cover scope, evidence, measures, security boundaries, deliverables, timeline, pricing and what may need a separate evaluation or implementation engagement.
Share your contact details and requirement. DataConsultant can review the likely assessment domains, evidence needs, stakeholder involvement and appropriate next step.