A 20 Question Scorecard for Your Live Sales AI Pilot
Evaluate live sales AI with a practical 20-question scorecard covering answer accuracy, source quality, and in-call usability before expanding your pilot.

A live sales AI pilot should end with a decision you can explain. Which questions can reps handle with it? Where do they still need a specialist? What would make you stop the rollout?
Those answers are easy to lose when a pilot becomes a series of impressive demos. A system might find the right document but omit the condition that matters to the buyer. It might generate a useful answer after the conversation has moved on. Reps might like it while quietly checking every response somewhere else.
Build the scorecard before the first test. Then use the same questions, evidence requirements, and failure rules throughout the pilot.
Define the pilot before scoring it
Choose a narrow sales motion, such as early technical discovery for one product. Include both newer and experienced reps. Assign a sales engineering or product owner to judge answer correctness, an enablement owner to review approved wording, and a RevOps owner to keep the test log.
Build a question set from recurring buyer questions your team is authorized to use. Include straightforward questions, follow-ups, ambiguous wording, and questions the selected knowledge does not answer. Remove unnecessary customer information. For each question, record the expected answer, essential qualifiers, acceptable source, and whether the correct response is to clarify or escalate.
Run internal rehearsals first. Use customer calls only after the relevant security, access, meeting-disclosure, and consent requirements have been addressed. The pilot itself does not waive those requirements.
Use a simple scoring rule
Score each of the 20 criteria from zero to two. Every criterion has the same weight:
- 0: Fails the agreed test, or there is no evidence.
- 1: Works inconsistently, needs a workaround, or requires a material correction.
- 2: Meets the requirement across the agreed test cases, with recorded evidence.
Each of the four categories has a maximum score of 10, for a total of 40. Record a score and supporting evidence for every criterion, plus an issue owner and retest result where needed. Keep category scores and failures visible. A high total cannot compensate for exposing restricted information or inventing a contractual commitment. Set those stop conditions with the relevant owners before testing.
Five questions about answer quality
1. Does the response answer the buyer's actual question? Test intent, not keyword matching. A question about migrating existing accounts should not pass because the tool explains how to create a new account.
2. Does it preserve the qualifiers? Check plan, region, deployment type, version, prerequisites, and exclusions. A short answer is useful only if shortening it keeps the meaning intact.
3. Does the cited source support the claim? Open the source and locate the relevant passage. A link to a related document is not enough to validate the answer.
4. Does it handle conflicting or superseded information appropriately? Include an older document and a current approved answer in a controlled test. Record which appears and whether the conflict is visible.
5. Does it avoid inventing an answer when the evidence is missing? Include unsupported questions. Check the actual output and teach reps what to do when support is absent; do not assume any system will abstain reliably.
Five questions about live use
6. Does the question get recognized in natural conversation? Include interruptions, paraphrases, and questions embedded in a longer statement. Record missed questions and irrelevant triggers separately.
7. Is the answer ready while it is still useful? Measure from the end of the spoken question to the first usable, supported answer. Keep slower cases visible rather than relying only on an average.
8. Can the rep understand the answer at a glance? Ask the rep to explain the answer in their own words. If they need to stop listening to decode it, record that friction.
9. Can the rep inspect the source without losing the conversation? Test the actual device and meeting setup. Record the steps needed to open evidence and return to the call.
10. Does a follow-up preserve context without importing assumptions? Ask a second question that changes a condition, such as the deployment model. The earlier answer should not carry over unexamined.
Five questions about knowledge and boundaries
11. Is the required knowledge actually available in this workflow? Verify the live answer surface separately from chat, search, or CRM features. An integration listed on a website does not establish that every workflow can use it.
12. Can the content owner correct an answer at its source? Make a controlled correction, follow the documented refresh process, and repeat the original question. Record the version the system uses.
13. Does the permission model fit your intended users? Test with designated accounts and harmless sample documents. Ask administrators to verify both ingestion scope and answer visibility before adding sensitive material.
14. Can the team distinguish reusable guidance from account-specific exceptions? A concession approved for one customer must not become the general answer for everyone else. Include a sanitized exception case in the test.
15. Are the supported platforms and operating conditions clear? Test the exact operating system, meeting application, audio configuration, language, and network conditions planned for the rollout. Record unsupported combinations explicitly.
Five questions about rollout readiness
16. Can reviewers reconstruct a failure? Keep the question, expected result, actual answer, source, relevant configuration, and reviewer decision. A screenshot alone may omit the context that explains the problem.
17. Do reps know when to verify, clarify, or escalate? Rehearse those choices. A rep should recognize when an answer crosses from documented capability into an implementation promise or an exception requiring approval.
18. Is there a named owner for each recurring failure? Assign knowledge defects to content owners, workflow issues to enablement, and product behavior to the vendor or technical owner. Give every issue a retest condition.
19. Does the pilot improve the chosen workflow against a baseline? Compare how reps handle similar questions with and without assistance. Record lookup effort, unresolved questions, and correction needs. Describe the comparison's limits.
20. Is the expansion decision specific? Write down which team, question types, knowledge sources, and meeting setups are approved for the next stage, plus the unresolved restrictions. Avoid a blanket approval based on a narrow pilot.
A hypothetical result worth catching
Imagine a buyer asks whether single sign-on is included in a proposed package. The system quickly answers yes and links to a product guide. The reviewer opens it and finds that the feature requires a different plan.
The answer arrives quickly but fails the qualifier test. Because question 7 measures time to a usable, supported answer, this response does not pass the live-use test either. The source link makes the mistake inspectable; it does not make the answer correct. Establish whether the problem sits in the source material, retrieval, or answer generation, then rerun the case with the corrected conditions.
This is a hypothetical evaluation example, not a Tenali customer result or a statement about Tenali pricing.
Make the decision from the failures as well as the score
At the final review, bring the scorecard, the unresolved issue log, and the conditions for expansion. Separate failures the team can fix by improving knowledge from behavior that still needs product work. Repeat failed cases after changes, and retain previously passed cases to check for regressions.
Use these decision bands as editorial starting points, not validated benchmarks or automatic rollout approvals. Agree on thresholds before testing and apply your stop conditions regardless of the total.
- 34–40: Consider a limited expansion within the tested scope if no stop-condition failures remain unresolved and remaining gaps have owners.
- 26–33: Keep the pilot contained. Fix the weak categories and retest before expanding.
- 0–25: Pause expansion. Narrow the use case or rebuild the knowledge and workflow before another evaluation.
A short pilot may show that reps can find supported answers with less interruption. It will rarely establish a causal change in win rate or sales-cycle length on its own. Keep the conclusion at the level the evidence supports.
Tenali surfaces answers from company knowledge with source links during live conversations. To evaluate that workflow for your team, book a demo and bring representative questions your team is permitted to share. Use the scorecard to make the next conversation concrete.
