Manager
AI Quality
How well CountPilot reads your reps — measured on the words they actually said.
Count lines read——
Right first time—matched, never corrected
Rep corrections——
Counts flagged—waiting on keep-or-fix
Golden set
—
Each case is the exact words a rep said and the lines they mean. A line counts only with the right SKU, count and location; wrong-SKU risk counts lines auto-matched to the wrong product — the number to keep at 0. Holdout cases are scored, never used to tune.
No runs yet — score the local parser to get a baseline.
Cases
What reps corrected
0What the reader heard vs. what the rep settled on. Add a correction to the golden set and every future change is scored on it.
No corrections yet — reps haven't had to fix a reading.
Reps disagree
—
No split votes — every phrase reps resolved means one thing.
Open Alias Review →AI spend this month
—
Match outcome
- Matched to a real SKU —
- Ambiguous · low-confidence —
- Unmatched product —
- Skipped —
Matcher owns SKU identity; ambiguous fragments stay in review until a rep confirms a real SKU.