# Manual acceptance scorecard

Run ID / date:
Product / version / model when disclosed:
Fixture version and engine:
Identity and granted data scope:
Schema comments / training snapshot:
First attempt or assisted retry:
Maximum permitted time / cost and how measured:

| Case | Expected | Actual rows | SQL executed? | Numeric pass? | Explanation pass? | Time / cost | Failure reason / correction |
| --- | --- | --- | --- | --- | --- | --- | --- |
| N1 | [[13500]] | | | | | | |
| N2 | [[3]] | | | | | | |
| N3 | [[1, 13500], [2, 0]] | | | | | | |
| N4 | [[7000]] | | | | | | |
| N5 | [[3000]] | | | | | | |
| N6 | [[40000]] | | | | | | |
| N7 | [[0]] | | | | | | |
| N8 | [[15000]] | | | | | | |

Development numerical result: __ / 2. Held-out-from-training numerical result: __ / 6. Overall: __ / 8 (descriptive only).

B1 clarification: not tested / pass / fail. Context and evidence:
B2 actual backend restriction: not tested / pass / fail. Identity, attempted query and returned-data evidence:

Do not combine behavior checks with numeric accuracy or average away an access failure. No cells are prefilled as product passes. The expected answers are reference fixtures, not observed AI answers.
