# AI pilot evaluation worksheet

Blank template · TheTechStack · https://www.thetechstack.com/guides/evaluate-an-ai-pilot

Use this alongside your pilot plan. Nothing in this blank worksheet is a completed test or release approval. Keep evidence in your team's approved systems; record references here rather than copying sensitive inputs or outputs.

## 1. Define the comparison before testing

- Workflow and task boundary:
- Pilot owner / reviewer / decision owner:
- Test dates:
- Current process (include existing tools and human review):
- AI-assisted process and human review required:
- Product/model version, configuration, enabled sources and actions:
- Population represented by these tasks; exclusions and known sample limits:
- How tasks were selected, including ordinary, difficult, and failure cases:
- How you will keep the two approaches comparable (same inputs, quality standard, reviewer, timing method; account for practice effects):
- Criteria agreed before testing, with thresholds and an owner for each:
- Mandatory stop conditions and who can pause access or actions:

Do not choose thresholds after seeing results without recording the change and rerunning the affected tests. Test an appropriate sample for your workflow; this template does not prescribe a universal sample size.

## 2. Plan the cases — no outcomes yet

Repeat this block for each case. Include normal work and cases involving missing/conflicting sources, denied or revoked access, unsupported requests, unavailable dependencies, and escalation or recovery where relevant. Use approved test identities and data.

- Case ID and category:
- Task and approved input reference:
- Expected acceptable result / scoring rule:
- Required source and access boundary:
- Expected refusal, approval, escalation, or recovery behavior:
- Risk if the case fails:
- Tester / reviewer:

## 3. Record observations — one block per case and attempt

Preserve failed attempts. Record a retest as a new attempt linked to the original, including what changed.

- Case ID / attempt / date:
- System configuration reference:
- Current-process result and evidence reference:
- AI-assisted result and evidence reference:
- Current-process active work minutes (including review and rework):
- AI-assisted active work minutes (including review and rework):
- Current-process end-to-end elapsed minutes (including waiting):
- AI-assisted end-to-end elapsed minutes (including waiting):
- Corrections required / reviewer disagreement and resolution:
- Quality criterion: pass / fail / not tested / inconclusive:
- Access and action boundaries: pass / fail / not tested / inconclusive:
- Escalation or recovery: pass / fail / not tested / inconclusive / not applicable (explain):
- Observed cost, units, and billing/evidence reference (or unknown):
- Stop condition triggered? Action taken, owner, and unresolved impact:
- Reviewer / review date:

A planned case is not a pass. Missing evidence stays not tested or inconclusive. Record denied access and blocked actions as observed behavior, not assumptions based on a vendor description.

## 4. Compare results at the same quality bar

- Cases planned / attempted / reviewed / excluded (and reasons):
- Quality passes / reviewed cases (show numerator and denominator):
- Boundary failures and unresolved stop conditions (list individually):
- Current-process total active minutes / AI-assisted total active minutes for the same reviewed cases:
- Active minutes saved = current-process total minus AI-assisted total:
- Elapsed-time difference, reported separately from active work:
- Unacceptable or incomplete results excluded from time-saving claims (list and explain):
- Estimated cost at expected volume:
- Assumptions: volume, licenses, usage, reviewer time/rate, setup, maintenance, support:
- One-time costs separated from recurring costs:
- What this sample cannot establish:

A faster unacceptable answer is not a successful task. Do not bury an access failure in an average quality score. Label cost projections as estimates and keep them separate from observed spend.

## 5. Make a bounded decision

- Decision: continue within the current pilot / fix and retest / stop / propose limited expansion:
- Evidence supporting the decision:
- Open failures, missing evidence, and required retests:
- Action / accountable owner / due date for each open item:
- Approved scope and remaining restrictions:
- Monitoring owner, review date, and pause/rollback procedure:
- Decision owner / decision date / approval reference:

Leave the decision pending until an accountable person reviews the evidence. A proposed expansion is not approval to expand. This worksheet supports a team decision; it does not certify security, compliance, or production readiness.
