A buyer’s checklist for a paid enterprise retrieval POC
A paid retrieval POC should test relevance, answer support and permissions against agreed criteria, with costs and exit terms set before a go/no-go decision.

Before you pay, agree on the decision, test and exit.
- Set a decision gate: Name the production choice the POC will inform, then define its scope, deliverables and pass conditions.
- Test distinct outcomes: Score retrieval, answer support, permissions and operating effort separately.
- Use buyer-controlled evidence: Freeze a representative corpus, user roles and difficult queries, and record expected evidence before the final run.
- Price the full boundary: Document the POC fee, usage assumptions, buyer effort, responsibilities, deployment, data handling and closeout terms.
A paid retrieval proof of concept (POC) should resolve a bounded technical and operational question: can the proposed service find the right evidence for specified user groups without exposing restricted material? Agree on the decision and the test before configuration begins. That gives both sides a shared basis for evaluating the fee and deciding what happens next.
1. Decide what the POC must prove
Start with the business decision, not a feature list. Write one sentence stating what the organization will decide at the end: for example, whether a retrieval system can support a named policy workflow for specified user groups and repositories. Say whether the result will support production design, a bounded follow-up pilot or a decision to stop.
The UK government’s Testing and Piloting Services guidance recommends clear objectives and success criteria, with a method for scoring results. It describes a POC as a test of whether a proposed solution can meet operational goals, while noting that a POC may not establish every process needed for a final service. Treat production architecture and operating design as separate questions if they need more evidence.
Before requesting a proposal, write a scope note covering:
- Decision and users: Name the decision owner, sponsor, technical evaluator and user roles. Specify what each role will try, such as finding the current policy or identifying the source behind an answer.
- Sources and data: List repositories, content types and approximate sample size. Identify source owners and content that is current, obsolete, duplicated, scanned, structured or restricted.
- Supplier and buyer work: State what the supplier will configure, connect, parse, index, tune and support. List buyer tasks such as providing accounts, approving access and labeling expected evidence.
- Test boundaries: Set dates, query set, identities, environment, permitted services and scorecard version. State whether tuning is allowed on a development set before a separate, fixed evaluation.
- Deliverables: Request a test plan, data-flow and configuration description, acceptance evidence, issue log, cost assumptions and findings tied to the agreed criteria.
- Decision rules: Set measurable targets and critical failure conditions before the test. Examples include exposing restricted content, treating an obsolete document as current or omitting a required source citation.
A two-page scope note can be a useful illustrative length, not a mandatory format or contract requirement. A specific UK Department for Work and Pensions proof-of-concept contract example includes scope, high-level architecture and exit criteria. Use it as an example of making scope and exit explicit, not as a universal rule.
Keep the scope narrow enough for a controlled comparison. Record additional repositories, user groups or agent features as later-stage options unless they are needed to answer the original question. Approve scope changes in writing and update the scorecard, effort and timeline before extra work begins.
2. Score retrieval, answer support and task outcomes separately
A useful scorecard separates whether relevant passages were found, whether the answer faithfully reflects those passages, whether the information is authoritative and current, whether access rules were respected, and whether the user completed the task. Fluent wording is not evidence of correctness. An answer can faithfully repeat an outdated or incorrect source, so faithfulness to retrieved text must not be treated as factual or business correctness.
Stanford’s Introduction to Information Retrieval defines precision as the share of retrieved documents that are relevant and recall as the share of relevant documents retrieved. For ranked results, evaluate the top-k positions users actually see. For example, if three of the first five passages are judged relevant, Precision@5 is 3/5, or 60%. This is an illustrative calculation, not a benchmark or recommended target.
Define the unit being scored and how duplicates are handled. For example, count unique source passages after a stated deduplication rule, rather than letting repeated chunks inflate a result. Full-corpus recall is meaningful only when the relevant-document set is sufficiently complete and labeled. If the team has labeled only the expected evidence for selected queries, call the measure required-evidence coverage, define its denominator, and do not present it as full-corpus recall. Record the relevance instructions and have human reviewers adjudicate disagreements or report agreement and unresolved cases.
Include unanswerable queries in the test and score correct abstention separately. Define whether they are excluded from retrieval metrics, how they affect answer/task scores and what counts as an appropriate no-answer response. The original RAGAS paper distinguishes context relevance, answer faithfulness to context and answer relevance to the question. Use these as separate evaluation concepts, alongside an explicit review of source authority, version and conflicts.
| Scorecard dimension | What to measure | Evidence to retain | Decision rule to set before the run |
|---|---|---|---|
| Retrieval relevance | Whether retrieved passages meet the stated information need, including top-k quality | Ranked passages, deduplication rule and human relevance judgments | Select a top-k precision or graded-relevance target for the use case |
| Required-evidence coverage | Whether the expected sources or passages are retrieved | Labeled expected evidence compared with retrieved source identifiers | Define the labeled denominator; call it coverage unless the relevant set is complete enough for recall |
| Answer support and source quality | Whether material claims follow from the cited text, and whether sources are authoritative, current and non-conflicting | Claim-to-passage review, source/version checks and unsupported-claim flags | Define how unsupported, stale, conflicting or low-authority evidence affects acceptance |
| Abstention | Whether the system declines when the corpus cannot support an answer | Unanswerable queries, response and reviewer judgment | Define acceptable abstention and any penalty for unsupported answers |
| Task success | Whether representative users complete the stated task accurately | User outcome, time or effort and completion record | Set task and usability criteria relevant to the workflow |
| Permission correctness | Whether each identity can access only permitted material throughout the response path | Role-by-query results, including denied and revoked cases | Treat defined restricted-data exposure as a hard failure |
| Performance and operations | Whether latency, errors, updates and operating effort fit the intended use | Timed runs, update tests, issue records and named owners | Set thresholds from the buyer’s workload and operational needs; there is no universal target |
Keep hard gates alongside any weighted score. A high average must not cancel a permission failure. Set metric weights from the consequences of the use case, and specify the scoring method, population, denominator, treatment of unanswerable questions, duplicate handling, ties, missing traces and invalid runs before testing.
If an LLM judge is used, record the judge model, prompt and version. Calibrate it against human-reviewed samples, retain those judgments and report disagreements. An automated score is an evaluation aid, not ground truth. Set thresholds from the buyer’s risk tolerance and baseline, not from a number proposed after the supplier has seen results.
3. Build and protect a representative test set
Build a manageable sample from the target corpus and label expected evidence before the supplier runs the system. The BEIR information-retrieval benchmark paper evaluates varied retrieval tasks, domains and text types, and reports that in-domain performance does not by itself predict zero-shot generalization. For a buyer, that supports testing the kinds of company content and questions in scope rather than relying on a polished demo set.
Cover four dimensions:
- Document types: Include formats used in the intended workflow, such as HTML, office files, PDFs, tables and scanned documents. Sample short and long items, mixed layouts and important identifiers.
- Content conditions: Include current, superseded, duplicated and near-duplicate material, with different update dates and owners. Identify which source or version is authoritative for each test.
- Question types: Include direct facts, natural-language questions, exact-name or code lookups, ambiguous and multi-part questions, questions requiring multiple sources, and questions the corpus cannot answer.
- User roles: Include identities with realistic differences in access, including permitted and prohibited access to selected documents and roles with overlapping permissions.
For each query, record its task, expected sources or passages, acceptable alternatives, answer limits, test identity and whether it should be unanswerable. Store the query set and labels under buyer-controlled access. Do not give suppliers access to held-back queries or labels before the scored run unless the agreed test explicitly requires it.
Separate development questions from the fixed evaluation set. Use the development set for parsing, ranking or prompt tuning. Before the final run, freeze the evaluation set, expected labels, corpus snapshot and system configuration. Agree on a bounded number of final runs, what changes trigger a retest, how retests are scored and who approves them. Record per supplier the tuning hours, manual fixes, support personnel and related spend so results reflect the effort required, not just the final score.
Preserve the same corpus snapshot, identities and relevant configuration for supplier comparisons. Record file versions, ingestion dates and permissions. Include a repeat run where useful, and investigate corpus or configuration changes before attributing a changed result to system quality. If the corpus is small, report the limited scope rather than implying that a handful of examples represents every document.
4. Test performance, data handling and permissions
Define performance conditions before testing. Record p50 and p95 response times, concurrency, warm- and cold-cache conditions, error and timeout rates, update lag, and representative document and query sizes. Record the environment and workload for each run. Set thresholds from the buyer’s use case; no single latency or error target is universal.
Translate security review into observable tests and written data-flow answers. Identify where source data, extracted text, embeddings, logs, backups and model requests are processed or stored, who can access them, and which external services or subprocessors receive information. Have security, privacy and source-system owners review the flow before sensitive content is loaded.
NIST SP 800-53 Revision 5 discusses audit-record content and privacy risks from audit trails. Set access and retention for evaluation logs under the organization’s privacy and records rules; logs can themselves contain sensitive information.
Create a permission test matrix. For each identity, list permitted and prohibited documents and queries. Check search results, generated answers, cached answers, session or conversation memory, snippets, citations and attachments—not only the index. Test new grants, revocations and role changes. Record the end-to-end cutoff time from a permission change until no restricted content can be returned through any tested surface; set an acceptable window from the buyer’s policy and use case.
Test the whole path after index or content updates. A denied document should not appear in results, answer context, citations, user-visible logs or attached output for an unauthorized identity. Preserve a safe trace of failures and follow the organization’s incident process if production data is exposed.
For scoped document prompt-injection tests, use approved synthetic data. Test whether instructions embedded in retrieved documents can override user or system rules, reveal restricted data or trigger unauthorized tools. Record the exact test and outcome. A successful test set is evidence about those cases only; it does not prove immunity to prompt injection.
The OWASP Top 10 for LLM Applications (2025) is a reference for considering risks such as sensitive-information disclosure and supply-chain exposure. Inventory the models, external APIs, software components and data services in the POC, along with applicable data-use terms. Use data approved for those routes and ask the supplier to demonstrate the configured flow rather than relying on general assurances.
Write down the POC environment and responsibilities: hosting arrangement, network routes, credentials, encryption and key ownership where relevant, administrative roles, support access, incident contacts, data retention, deletion and deletion verification. NIST’s media-sanitization discussion supports documenting sanitization and disposal actions and verifying them before disposal. Agree on a closeout record showing what was deleted, retained under an approved basis or returned.
5. Define query-level evidence and safe trace handling
Agree on the evidence needed to judge acceptance, then separate it from optional diagnostics. For each acceptance query, request a reviewable record connecting the query and test identity to retrieved evidence, answer and outcome. Request diagnostic fields only where the system exposes them and they are within the agreed scope.
Acceptance evidence may include:
- Exact query, test-role identifier and timestamp, plus any query rewrite needed to explain the result.
- Retrieved document identifiers and versions, reviewable passage excerpts, rank and applied filters.
- Access decision, final answer, citations and the source text each citation points to.
- Configuration, model or index version needed to interpret or reproduce the test, plus errors, latency and manual intervention.
Do not request hidden chain-of-thought. Agree instead on observable evidence such as retrieved passages, citations, access decisions and the final response. Define how traces will be redacted, protected, accessed and retained; minimize personal or sensitive data and include deletion or retention in closeout.
The RAGAS evaluation framework distinguishes context relevance, faithfulness and answer relevance. Use query-level evidence to determine whether a failure came from a missing source, parsing, filtering, ranking, an unsupported claim or an incorrect citation. Ask what a supplier’s internal score means before comparing it; use it for troubleshooting unless a common interpretation has been validated.
Keep a concise failure log with category, example, impact, owner and retest result. Useful categories include wrong or stale source, missing or irrelevant passage, permission error, unsupported answer, broken citation, slow response, failed update and prompt-injection behavior. If configuration changes, rerun the original query and relevant neighboring cases against the same corpus snapshot, following the agreed retest rules.
6. Verify integrations, deployment and operating ownership
Inventory each source and system the POC touches. Ask the supplier to classify work as standard configuration, custom integration or buyer-owned setup. For each connector, check supported content and metadata, permission handling, initial and incremental loading, edit and deletion propagation, and how errors reach administrators. Test with controlled additions, edits and removals.
Separate the retrieval pipeline from agent operation. If the POC includes agents or tools, test tool discovery, authentication, allowed actions, input handling, approval controls and the audit record for each tool call. A managed agent does not automatically provide governance: verify the actual authorization, audit and approval controls in the tested configuration.
Seahorse Cloud’s product overview describes an integrated document-to-agent stack: object storage, a vector database, document parsing, semantic chunking and automatic vector database synchronization, alongside managed agents with MCP-standard tool calling. It lists on-premises installation and SaaS subscription as deployment options. This makes it relevant to evaluating an end-to-end path from document ingestion to agent tool use. A POC should still test permission propagation, recovery behavior, audit records and approval controls in the selected configuration; the listed capabilities do not establish those controls.
Document deployment location separately from operating responsibility. Ask who manages updates, backups, support access, connector configuration, source credentials, parser corrections, index monitoring, agent releases, incident response and production migration. Capture handoff artifacts such as architecture and data-flow diagrams, configuration exports where available, query sets, traces, known limits, runbooks and named operational owners. Count buyer setup and correction time as operating effort, not invisible POC work.
7. Put fees, operating costs and exit terms in writing
State what the POC fee buys: scope, dates, supplier roles and hours, response expectations, included usage and separately billed work. Keep the POC price distinct from production estimates. The U.S. federal FAR 8.405-2 describes statement-of-work elements for the federal ordering procedure it covers, including work, performance period, deliverables and standards. Commercial buyers can use similar categories as a drafting checklist and adapt them with procurement and legal advisers.
Use a worksheet to estimate the full POC cost:
Total POC cost = POC fee + extra usage + buyer hours × loaded hourly rate + separately billed rework and exit costs.
Do not count labor or usage twice if it is already included in the fee or another line. Keep any production estimate separate and state its usage assumptions. For example, a hypothetical POC with an $8,000 fee, $400 of extra usage, 30 buyer hours at $60 per hour ($1,800), and $500 in separately billed rework or exit costs totals $10,700. These invented figures demonstrate the calculation only; they are not a benchmark or a Seahorse Cloud price.
| Cost area | Questions to document |
|---|---|
| Setup and services | Which connectors, parsing, tuning, integration and support hours are included? What triggers paid change requests or rework? |
| Data and storage | What are the ingestion, storage, indexing, backup and transfer charges, where applicable, and what usage assumptions apply? |
| Retrieval and models | What query, embedding, reranking, inference or tool usage is included? What rates or tiers apply above the allowance? |
| Buyer effort | Which buyer roles and hours are expected? Which effort is already included elsewhere in the worksheet? |
| Production transition | Which configurations, evaluation artifacts and integrations may carry forward, and which work may need to be repeated or re-priced? |
| Support | Which hours, channels, response targets and supplier roles apply during and after the POC? |
Record currency, taxes, payment milestones, renewal and default-billing status, any credit toward a later purchase and its conditions. Set a spend cap and require approval before overages or added work. Ask how quoted usage maps to expected query volume, document growth and update frequency; label production scenarios and identify any binding amount.
Set an end date and decision breakpoint. The UK government’s Digital, Data and Technology Playbook warns about pilot creep and recommends contractual breakpoints at the end of stages. Specify whether extension requires written approval, a new price and a defined additional test. State who accepts deliverables and how incomplete or failed milestones are handled.
Closeout terms should cover return or deletion of source files, extracted text, embeddings, query logs, backups and test accounts; evidence of deletion; and any backup retention approved by the buyer. Include shutdown of paid resources, credential revocation, access termination, data export and handoff. Record rights to retain test results and configuration artifacts, and ownership and permitted use of custom code, schemas, prompts and other work products. Verify closeout rather than assuming it happened.
8. Make the go/no-go readout evidence-based
The final readout should compare observed results with the criteria approved before kickoff. Ask buyer evaluators and the supplier to provide the scorecard, corpus snapshot, test identities, query-level acceptance evidence, pass and failure counts, operating-effort record, cost assumptions and unresolved issues. Link conclusions to evidence and state which findings apply only to the tested corpus, environment and roles.
The UK government’s pilot-testing guidance calls for findings to be compared with success criteria, performance gaps to be identified and results to inform the next procurement stage.
Use a three-part recommendation:
- Proceed to the next stage when every critical gate passes and measured outcomes meet the agreed requirements. Name remaining production questions—such as scale, availability, broader integrations or support—and assign owners and decision dates.
- Run a bounded follow-up when one specific gap could change the decision. Define the question, added scope, evidence, fee, end date and revised exit condition before extending.
- Stop or select another option when a critical gate fails, retrieval falls below the agreed threshold, or the operating model and commercial terms do not fit. Preserve the failure evidence and decision rationale.
Record accepted and incomplete deliverables, any remedy or follow-up, final data and credential disposition, paid-resource shutdown and owners for next steps. A POC result does not establish production readiness beyond the tested scope. If proceeding, convert confirmed assumptions into production requirements and test them in the target environment. If stopping, close access and verify data disposition under the agreed terms.
How many queries should a paid retrieval POC include?
Include enough queries to evaluate each pre-agreed criterion across selected document types, roles and tasks. State that results apply only to the tested scope, and report the labeled population and its limits.
Should a supplier tune the system on our test questions?
Use a development set for tuning and reserve a separate, buyer-controlled set for the scored evaluation. Freeze its queries, labels, corpus and configuration before the bounded final run. Record tuning effort and follow the agreed retest rules.
Should retrieval and generated answers get one combined score?
No. Score retrieval, source quality, answer faithfulness, task success and permissions separately. A weighted score must not conceal a critical access failure, and a faithful answer can still repeat a wrong or outdated source.
What makes a paid POC worth extending?
A specific unresolved question that could change the decision may justify a bounded extension. Agree on the test, cost, owner, end date and exit condition before extending; do not let the pilot continue by default.
Evaluate your retrieval use case
Put these criteria to work against your own enterprise retrieval workflow.
