TL;DR
- Judge observed performance, not the length of a feature list.
- Score eight criteria from 0 to 2: 0 means missing, 1 means partial or manual, and 2 means proven end to end.
- Use 12 out of 16 as a starting pass mark, then adjust it to your risks.
- Reject any candidate that invents an answer during the unknown test or fails a required human handoff.
- Run the same lead, support, routing, and escalation tests on every candidate.
A prospective customer asks about price, agrees to leave contact details, and then asks a policy question your website does not answer. That third exchange reveals the real buying risk. An AI chatbot for your website may handle the easy question yet guess at the missing answer, lose the lead, or send the visitor to a person without the conversation history. The eight-part scorecard below gives you a repeatable way to test reliable answers, lead conversion, and safe handoff before you commit.
Key Takeaways
- A reliable chatbot must recognize when approved information cannot support an answer.
- Lead capture deserves a full score only when usable details and conversation context reach the correct destination.
- A human handoff must preserve enough context for the visitor to continue without starting over.
- Content coverage and freshness need separate tests because a once-correct answer can become outdated.
- Phone, multilingual support, and advanced actions should affect the result only when the business actually needs them.
Set the pass mark before you test a website chatbot
Define the scoring rules before you start a trial. Otherwise, an attractive demo or one impressive answer can change how you judge later candidates.

Score each of the eight criteria on the same scale:
- 0: Missing. The required behavior is unavailable or fails during testing.
- 1: Partial. The behavior needs manual work, loses useful information, or remains unverified.
- 2: Proven end to end. You observe the complete visitor and staff path working as required.
The maximum is 16 points. Use 12 as a practical starting pass mark, not as a validated industry benchmark. A business answering routine opening-hours questions may accept that threshold. A team handling account disputes or sensitive information may require full scores on grounding, access controls, and escalation.
Add two failure gates that override the total:
- Reject a candidate that confidently invents an answer during the unknown test. A safe response should acknowledge the limit, ask for clarification, refer to an approved source, or offer an appropriate handoff.
- Reject a candidate that cannot complete a mandatory human handoff. Displaying an email address does not count if your process requires routing, notification, or preservation of the conversation.
For example, six full passes and two partial passes produce 14 points. That result clears the starting threshold, but the candidate still fails if a partial score hides a broken support escalation that your business requires.
Equal weighting makes the first comparison simple. Change the weighting when your workflow justifies it. A business losing after-hours inquiries may prioritize lead capture and routing. A company with frequently changing policies may require full scores for grounding and freshness.
Score the eight jobs that determine real visitor performance
Use the Eight-Part Website Chatbot Scorecard for every candidate. Save an answer, transcript, notification, destination record, or screenshot beside each score. A feature-page claim is not enough for a full pass.
| Criterion | Observable test | Full-pass condition |
|---|---|---|
| 1. Answer grounding and citations | Ask one question covered by an approved page and one absent from every approved source. | The chatbot answers from approved information, identifies its source when configured to do so, and handles the unknown safely. |
| 2. Content freshness | Change a test price, date, or policy in a connected source, then check the answer again. | A defined update process produces the corrected answer within a period your business accepts. |
| 3. Lead capture | Offer a name and fictional contact method after asking about price or availability. | The required details and conversation context reach the intended person or system. |
| 4. Escalation and human handoff | Ask for a person after an unsupported or account-specific question. | The request reaches the correct destination with enough context for the conversation to continue. |
| 5. No-code setup effort | Connect a representative source and publish a private test version. | The team responsible for the chatbot can configure, test, and maintain it with available skills and access. |
| 6. Workflow and integration fit | Complete one required action, such as booking, quote routing, or support delivery. | The action finishes in the correct destination with the required fields and context. |
| 7. Conversation review and control | Find the test conversation, inspect its history, and record the outcome. | Staff can locate the exchange, understand what happened, identify failures, and track follow-up. |
| 8. Channel and brand fit | Check the experience on the devices, languages, and channels your visitors use. | The chatbot meets your defined brand, device, access, and channel requirements. |
Give a score of 1 when a path works only after someone copies information between systems, repairs missing context, or completes an unverified step. Give a 2 only after the intended result arrives where staff need it.
Keep conditional features out of the main decision unless they serve a real visitor path. A web-only business should not reward phone support at the expense of answer quality. Broad language coverage matters only when customers need those languages and the team can review the resulting conversations.
Run one trial script on every shortlisted chatbot
A fixed test sequence prevents one product from receiving easy questions while another faces difficult requests. Base the sequence on recurring visitor conversations and current public business information. Use fictional contact details, not real customer records.

Run these seven steps:
- Ask a pricing or availability question. Choose one with a clear answer on an approved page, such as, “What does the standard cleaning visit include?”
- Ask a policy question. Test a condition visitors often misunderstand, such as, “How much notice do I need to reschedule?”
- Ask an unknown question. Use a detail missing from all approved sources, such as, “Can your technician repair a rare imported appliance during the visit?”
- Start a lead path. Express interest and provide fictional contact details. Check what the chatbot collects and whether the request for consent fits the intended use.
- Request the next action. Ask to book, request a quote, or send the inquiry to the appropriate team.
- Request a person. Do this after the unsupported question so the escalation test includes useful context.
- Inspect the staff side. Find the transcript, notification, lead record, booking, or routed request. Verify its destination, fields, timing, and conversation history.
Record the exact wording and result at each step. If a person must repair the workflow, score it as partial unless manual handling is part of your intended process.
Test desktop and mobile when both matter to your traffic. Rephrase the unknown question at least once because a single safe refusal does not establish consistent behavior.
Five to seven questions can cover an initial pass, but that range does not prove reliability. Add cases until you have tested every high-frequency and high-consequence path your chatbot must handle.
Worked example: a local service business tests lead and support paths
Consider an illustrative home-maintenance company. The owner wants to capture inquiries after hours and answer routine appointment questions during the day.
The lead path begins with, “How much does a standard service visit cost, and do you have availability next Tuesday?” The chatbot should rely on approved pricing and scheduling information, collect fictional contact details, and route the request for a quote or booking.
The support path begins with, “Can I reschedule an existing appointment without a fee?” The visitor then asks about an unusual repair that the company’s pages never mention. The owner expects the chatbot to avoid guessing and offer a human handoff with the earlier messages attached.
The owner scores both candidates after running the same script and saving observed evidence:
| Criterion | Candidate A | Candidate B | Observed evidence |
|---|---|---|---|
| Answer grounding and citations | 2 / 2 | 2 / 2 | Both answered the approved pricing question from the supplied service page. |
| Content freshness | 2 / 2 | 1 / 2 | A reflected the revised fee within the accepted window; B still showed the earlier fee. |
| Lead capture | 2 / 2 | 2 / 2 | Both delivered the fictional name, contact method, and conversation context to the test inbox. |
| Escalation and human handoff | 1 / 2 | 2 / 2 | A routed without the earlier messages; B delivered the complete transcript. |
| No-code setup effort | 2 / 2 | 1 / 2 | A was configured by the owner; B required vendor help for one routing rule. |
| Workflow and integration fit | 2 / 2 | 2 / 2 | Both completed the quote-routing test with the required fields. |
| Conversation review and control | 2 / 2 | 1 / 2 | A exposed transcript and outcome filters; B exposed the transcript but no outcome label. |
| Channel and brand fit | 1 / 2 | 2 / 2 | A needed a mobile spacing fix; B passed mobile and brand checks. |
| Total | 14 / 16 | 13 / 16 | A fails the unknown-answer gate; B passes both gates. |
Imagine Candidate A scores 14 but invents a confident answer about the unusual repair. Candidate B scores 13, states that it lacks supporting information, and routes the complete exchange correctly. Candidate B stays on the shortlist. Candidate A fails because its higher average hides a prohibited behavior.
InsertChat can be tested as one candidate under the same rules. Its product information describes website-based setup, answers drawn from business content, source citations, lead capture, booking, routing, branding, and conversation controls. Treat these as capabilities to verify in your configuration, not as automatic passing scores. You can see how InsertChat works before running the same controlled test.
Know when an AI chatbot is more system than you need
Choose a simple FAQ widget when a small set of stable, low-risk answers covers nearly every visitor need. A fixed interface may be easier to maintain if five questions rarely change and no follow-up action is required.

Choose live chat when immediate human judgment is central. Disputed charges, unusual technical failures, and detailed custom estimates may move to a person so quickly that automating the opening exchange adds little value.
Use a human-first workflow for sensitive, restricted, disputed, or high-consequence cases. A chatbot may collect a basic reason for contact, but a qualified person should make the decision. Review privacy, security, access, and retention requirements before allowing real visitor data into any system.
As automation expands, tighten controls around approved sources, permitted actions, conversation review, and escalation. A chatbot that reports public opening hours needs fewer safeguards than one that changes appointments or retrieves account-specific details.
Ownership matters too. If nobody reviews failed handoffs, updates source material, or checks conversation quality, performance can decline as business information changes.
Turn the scores into a shortlist decision
Start with the failure gates. Remove any candidate that invents an answer in a critical test or cannot complete a mandatory handoff. Price, design, and optional channels should not rescue it.
Compare the remaining candidates on the paths tied to your main problem. For after-hours lead capture, focus on grounding, lead delivery, routing, and review. For routine support, give greater weight to freshness, safe escalation, and the team’s ability to identify unanswered questions.
Retest every score of 1 after configuration. A partial result may come from a missing setting rather than a product limit. Repeat the same test, confirm the destination and context, then update the score only if the complete path works.
Apply optional checks last. Phone, additional languages, rich displays, and advanced actions belong in the decision only when they support a defined customer journey.
Write down five to seven real visitor questions, run the same sequence on each candidate, and save evidence beside every score. Product-specific embedding, account, and connection questions can be checked in the InsertChat setup and product FAQs if it remains on your shortlist.
FAQ
How many questions should a chatbot trial include?
Start with enough cases to cover a common pricing or product question, a policy question, an unsupported request, lead capture, the required next action, and human escalation. Five to seven can cover the first pass, but no universal sample size proves reliability. Add variations for every high-volume or high-risk path before launch.
Should price determine the shortlist?
Compare price after each candidate completes the required answer, lead, and handoff paths. Include the staff time needed for setup, maintenance, manual routing, and conversation review. A lower subscription cost may bring greater operating effort when important steps remain manual.
Do source citations guarantee a correct answer?
No. A citation helps a visitor or reviewer inspect the basis of an answer, but the source may be outdated, irrelevant, or interpreted incorrectly. Check whether the response matches an approved source and what happens when no supporting information exists.
When should phone or multilingual support affect the score?
Include phone when missed calls are part of the business problem. Include additional languages when customers need service in those languages and the team can assess answer quality and escalation. Otherwise, treat both as optional requirements.
How often should the team review conversations?
Set the cadence according to traffic, risk, and the rate of business change. Review more often during launch and after changes to prices, policies, sources, or routing. Track unsupported questions, incorrect answers, abandoned lead paths, and failed handoffs. Then correct the relevant source or workflow and repeat the original test.



