AI Proof of Concept: What It Should Prove Before You Build
Last updated: September 2026. External figures are attributed and dated in the text. The sample-size table is illustrative arithmetic, not measured results.
In short: An AI proof of concept (POC) is a time-boxed build that answers one question on your own data: does this approach clear a quality, cost and latency bar that was agreed in writing before the work started? It does not prove the system is ready for production, and it should be allowed to end in "do not build." BCG's January 2025 AI Radar survey of 1,803 executives found 60% of companies failing to define and monitor any financial KPIs for AI value creation. In our reading, a POC without a written bar repeats that gap in miniature. Our POC in 3 Weeks, from $15,000, fixes the bar in Week 1 and ends in a build, refine or do-not-build verdict.
What is an AI proof of concept?
An AI proof of concept is a scoped, time-boxed build that tests whether one specific AI approach meets a pre-agreed bar on real data, at a cost and speed the business can accept. Its output is a decision, not a product. The four terms below are often used loosely; these are our working definitions.
POC, prototype, pilot and production compared
| Question it answers | Data | Who uses it | Success bar | Typical output | |
|---|---|---|---|---|---|
| Prototype | Can we show the idea? | Sample data | An internal audience | It looks convincing | A demo |
| Proof of concept | On our real data, does it clear the bar we set, at an acceptable cost and latency? | A representative real-data extract | The evaluators | Thresholds agreed in writing before the build | An evaluation report and a build, refine or do-not-build verdict |
| Pilot | Does it work for real users in a real workflow? | Live data, limited scope | A small group of real users | A business outcome for that group | Evidence for a wider rollout |
| Production | Does it keep working, and can we prove it? | Live data at full scale | All intended users | Outcome monitoring with a named owner | A running, verified system |
What should an AI proof of concept prove?
It should prove four things, each against a threshold set in advance: quality on your real data, cost per task at production volume, latency at a stated percentile, and reliability on edge cases and on the "I don't know" path. Quality alone is not enough. The Holistic Evaluation of Language Models study (Liang et al., arXiv:2211.09110, 2022) measured seven metrics across sixteen core scenarios instead of accuracy alone, and reported that before it, models had on average been evaluated on just 17.9% of those scenarios.
Two conditions come before the four measures. The task has to be technically feasible: Andrew Ng's AI Transformation Playbook (Landing AI, first released in 2018) advises having experienced AI engineers check a project's feasibility before kickoff. And the objective has to be clearly defined and measurable, creating business value that someone can point to. A POC built on a task nobody can score is a demo with a longer schedule.
Why write the pass/fail bar down before the build starts?
Because a bar set after the result is a rationalization. In its January 2025 AI Radar survey of 1,803 C-level executives across 19 markets and 12 industries, BCG found that 60% of companies were failing to define and monitor any financial KPIs for AI value creation (BCG, 2025-01-15). That survey is about AI programs, not proofs of concept. Our reading is that a POC which begins without a written bar is the same gap in miniature: nothing exists to grade the result against, so it is graded on impression.
In NewGenApps' POC in 3 Weeks, the thresholds for quality, cost, latency and reliability are signed off at the end of Week 1, together with the frozen test set. They change afterward only through change control, with the effect on the timeline stated first. That closes the door on a POC being quietly graded as a win whatever the result.
Gartner predicted in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, naming poor data quality, inadequate risk controls, escalating costs and unclear business value. That is a forecast, not a measured rate. Whatever the true rate turns out to be, the choice is when you find out: early, at a fixed scope, or late.
How big does the test set need to be?
Big enough that the bar can be decided by it. A pass rate measured on a small set comes with a wide range of plausible true values, and if your threshold sits inside that range, the set cannot settle the question.
Sample size and the width of the answer
Illustrative arithmetic, not measured results: the 95% Wilson interval for an observed 90% pass rate.
| Test cases | Observed passes | 95% interval for the true pass rate |
|---|---|---|
| 30 | 27 | 74.4% to 96.5% |
| 50 | 45 | 78.6% to 95.7% |
| 100 | 90 | 82.6% to 94.5% |
| 200 | 180 | 85.1% to 93.4% |
| 500 | 450 | 87.1% to 92.3% |
Read the second row this way: with 50 cases, a system that scored 90% could truly sit near 79% or near 96%, so a bar of "at least 90%" cannot be decided from that set. With 200 cases the lower end is about 85%, and with 500 about 87%.
The Wilson interval is one of the two that Brown, Cai and DasGupta recommend for small samples (Statistical Science, 2001); the simple normal-approximation interval is erratic at small sizes such as the first row. Two things make real uncertainty wider than this table: cases that are not independent, such as many rows drawn from one document or one customer, and errors in the labels themselves. Treat the table as the narrowest plausible range. In practice, set the size of the test set in Week 1 from the tolerance the decision can bear, draw it from real data including the awkward cases, and freeze it.
How do you scope an AI proof of concept?
Seven steps, in order:
- Pick one task. Narrow enough to evaluate unambiguously, such as classifying an inbound document or extracting fields from a contract.
- Name the decision the output informs. Who uses it, in what workflow, and what changes as a result. This sets whether the bar is "good enough to act on directly" or "good enough to speed up a human who stays in the loop."
- Write the thresholds. A target and a minimum for quality, a ceiling for cost per task and per month, a maximum p95 latency, and a maximum rate of confident-but-wrong answers.
- Split your real data. Freeze a test set from it, awkward cases included and sized against step 3. The frozen set is never used to steer changes. Keep the rest as a development slice.
- Build the evaluation harness with the system, and iterate on the development slice. Every meaningful change is scored on the slice, so progress is measured rather than asserted. The frozen test set is not touched.
- Score the frozen test set exactly once, on a frozen build, with someone other than the builder running it.
- Decide by the rule you wrote in step 3. Build, refine or do not build.
What does a 3-week AI POC look like?
The POC in 3 Weeks compresses define, build and verify into 21 days, on the client's own data.
The three weeks at a glance
| Week | Focus | Gate to advance |
|---|---|---|
| 1. Scope and design | Fix the use case and the input and output contract; sign off the thresholds; take delivery of the data extract; freeze the test set and set aside a development slice; design the approach | Signed-off metrics, a frozen test set, a development slice, and a reviewed design with its risks logged |
| 2. Build with the harness | Build on real data; build the evaluation harness in parallel from day one; score every meaningful change on the development slice, never on the frozen set | A working system producing scored results on the development slice, trending against the thresholds |
| 3. Evaluate and decide | Freeze the build; score the frozen test set exactly once, run independently; stress the edge cases and the "I don't know" path; write the report and the recommendation | An independently verified result on the running system, and a recommendation the budget owner can act on |
You receive a working proof of concept on your real data, an evaluation report, the evaluation harness, the codebase and an architecture document, and an honest recommendation. When the verdict is build or refine, a costed production roadmap is included. The full terms are on the engagements page.
What does a POC not prove?
A pass means "worth building," not "ready to run." Outside a POC's scope: production hardening (availability, scaling, full security review); production integrations with live systems, write-back and access control at scale; data engineering at scale, because the POC uses a representative extract; more than one use case; interface work beyond what the evaluation needs; model fine-tuning unless the agreed metrics cannot be met without it; and change management.
Those belong to a production build, and the questions to ask before shipping are in the AI production readiness checklist. A system can also pass every check and still deliver nothing, which is the point of liveness versus outcome.
When is "do not build" the right answer?
When the evidence says so: the approach does not clear its minimum bar on real data; the cost or latency economics fail at production volume; or the reliability profile is unacceptable for the decision the output informs. In those cases a reasoned "do not build," ideally with a different approach worth testing, is a successful outcome of the POC. Finding that out in three weeks at a fixed scope is cheaper than finding it out six months into a production program.
What does an AI POC cost?
The POC in 3 Weeks starts from $15,000 and runs for 3 weeks (21 days). The fee is set against a scoped engagement after a free 30-minute call. What drives an AI engagement's price, and how to compare quotes, is in how much AI consulting costs; the full ladder is on the engagements page.
Related reading and sources
To see what your own bar and test set would look like, book a 30-minute working session. Related reading: the evaluation harness, how you know an AI system works in production, deploy and verify AI, the POC-to-production rate to ask a vendor for, and AI Rescue for a pilot that has already stalled.
Sources: BCG, "One Third of Companies Plan to Spend More than $25 Million On AI in 2025 Amid Widespread Optimism for Autonomous Agents" (AI Radar 2025, n=1,803; press release, 2025-01-15); Landing AI, "AI Transformation Playbook" (Andrew Ng; first released December 2018); Gartner, "Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025" (press release, July 2024); P. Liang et al., "Holistic Evaluation of Language Models," arXiv:2211.09110 (2022-11-16); L. D. Brown, T. T. Cai and A. DasGupta, "Interval Estimation for a Binomial Proportion," Statistical Science 16(2): 101-133 (May 2001), DOI 10.1214/ss/1009213286; NewGenApps, POC in 3 Weeks scope and the engagements page (2026-09-29). Sample-size table computed by NewGenApps with the Wilson score formula, 2026-09-29.