Why test cases matter before rollout
A demo prompt can look useful with one clean example. Real workflows see incomplete notes, contradictory customer requests, missing prices, old policy snippets, private data, angry messages, after-hours requests, and staff shortcuts. This worksheet helps owners prove an AI workflow can handle normal work, fail safely, and hand off risky cases before it becomes part of the daily process.
Test-case worksheet
| Field | What to capture | Owner review rule |
|---|---|---|
| Workflow purpose | Task name, user, tool, connected systems, and the exact business outcome the workflow should support. | Approve one narrow workflow at a time; do not approve vague “AI assistant” behavior. |
| Normal cases | Three to five realistic examples from recent work, with approved source facts and expected output. | Output must match source facts and be useful without inventing missing details. |
| Edge cases | Missing information, angry customer, safety issue, legal/financial/privacy detail, out-of-policy request, and conflicting source notes. | Edge cases should trigger a hold, escalation, or draft-only response. |
| Forbidden outcomes | Claims, prices, discounts, warranties, availability, diagnoses, legal/tax/medical advice, credentials, approvals, refunds, or customer-data uses the AI may not invent. | If the AI invents or implies one, the workflow fails the test. |
| Review owner | Named person who checks outputs, signs off on rollout, owns rollback, and sets the next review date. | No owner means no rollout. |
Copy/paste test log
AI workflow test case log
Workflow: [name]
Tool/prompt/version: [tool + version]
Owner/reviewer: [person]
Allowed source facts: [approved notes, policy, CRM fields, web page, transcript]
Blocked data/outcomes: [private data, prices, warranties, legal advice, unsupported promises]
Test case type: [normal / edge / stop-rule]
Input sample: [paste realistic source notes]
Expected output: [what a safe answer should do]
Actual AI output: [paste result]
Pass/fail/hold: [decision]
Fix needed: [prompt, data, workflow, owner review, or stop]
Next review date: [date]
AI test-review prompt
Act as a cautious small-business AI workflow QA reviewer. Use only the test cases and verified source facts below. Do not invent approval history, policy, pricing, warranty, availability, credentials, refunds, diagnoses, legal/tax/medical advice, customer consent, or private facts. Return: 1) tests that passed, 2) tests that failed, 3) missing source facts, 4) stop-rule violations, 5) prompt/workflow fixes, and 6) whether this workflow is safe for draft-only, human-reviewed use, or no rollout.
Workflow:
Allowed source facts:
Blocked outcomes:
Test cases and AI outputs:
Owner/reviewer:
Rollout decision needed:Fast QA before rollout
- Include at least one case that should produce “I need a human to review this.”
- Test with messy real-world notes, not only polished sample inputs.
- Keep customer data, passwords, payment details, employee records, and sensitive private facts out of generic tools unless approved.
- Set the workflow to draft-only until repeated tests pass and a human owner signs off.
Related free assets: AI Output Acceptance Checklist, AI Draft Approval Queue, and AI Workflow Weekly Review Agenda.
Disclosure: Horizon Flow is Andrew Burton's digital product catalog. This worksheet is useful without purchase; product links are labeled and UTM-tagged.