Vision: Help visitors ask with images when words
Use owned content to answer visitor questions with less friction.
7-day free trial
What this feature covers
Why it matters
The practical reason to use it.
Vision helps when a visitor's question depends on what they can show.
How it works
A step-by-step look at the workflow.
Step 1
Choose the visitor journeys where images reduce back-and-forth, such as visual support, product identification, document help, or troubleshooting.
Step 2
Enable uploads for the relevant assistant and define what it may explain from the image versus what must come from approved sources.
Step 3
Let visitors upload a screenshot, photo, or document image during the conversation and ask their question in plain language.
Step 4
Review visual conversations, tighten prompts, add missing source material, and route cases to a person when the image needs human judgment.
Core job
The main job this feature handles.
Image understanding
Use screenshots, product photos, labels, and visual examples to reduce visitor explanation burden.
Document image help
Help visitors interpret forms, receipts, policy pages, or scanned resources when approved guidance exists.
Assistant-level control
Limit visual input to the assistants and pages where images improve support or discovery.
Sensitive-case routing
Route unclear or high-risk visual cases to a person instead of forcing an automated answer.
Daily use
How teams use it after launch.
Start with one bounded workflow
Use Vision on the narrowest workflow where the team can measure whether the feature reduces friction, improves clarity, and creates faster support.
Keep edge cases visible
Review the conversations, prompts, and system actions tied to vision so operators can see where the rollout still depends on manual judgment.
Connect the surrounding systems
Vision is stronger when the feature sits beside the knowledge, integrations, and routing rules that already determine what happens after the first.
Expand after the first proof
Once the first deployment is stable, teams can extend vision into more surfaces and assistants without rebuilding the same control model from.
Control points
What to keep controlled.
Review real visitor conversations
Use real conversation data to inspect whether vision is actually improving answer quality, reducing back-and-forth, and creating better product and resource guidance.
Check ownership and controls
Look at which team owns the feature, where approvals still matter, and how the capability interacts with surrounding systems.
Track downstream changes
A strong rollout shows up after the first response too: cleaner handoff, clearer escalation, less manual cleanup, and faster next-step execution.
Expand with evidence
Only widen the rollout after the first bounded workflow is clearly stable.
What you get
The changes teams should notice first.
- Faster support when screenshots or photos explain the issue
- Better product and resource guidance when images clarify context
- More complete handoff because uploaded context stays with the thread
- Tighter control by scoping uploads to approved visitor journeys
The facts do the selling
Plan facts, platform capabilities, and worked examples — every claim here is checkable, not a pitch.
White-label included — never a paid add-on. Copyright removal from $98/mo. Full white-label — custom domain, branded portal, your-domain emails — from $198/mo.
The white-label wedge
Platform fact
Training runs on your sitemap, PDFs, docs, and YouTube transcripts. Answers cite the source pages they came from.
Trained on your content
Platform fact
Five clients at $300/mo on a $198/mo Agency plan is $1,300+ of monthly margin before usage.
A 5-client agency on one flat plan
Worked example
Your questions, answered.
Tap any question about the product, pricing, security, or setup to see a straight answer.
InsertChat
Answers about InsertChat
Hi! Tap any question below and I'll answer it for you.
Vision questions
What kinds of questions fit vision?
Vision fits questions where a screenshot, product photo, label, receipt, form, or document image carries important context. It works best when your team has approved sources that explain what the assistant should do with that visual detail. The operational question is whether vision makes the workflow clearer once real conversations, real ownership, and real edge cases show up. That is the bar teams should use before they expand the rollout across more assistants, more channels, or more teams.
Can I keep vision limited to specific assistants?
Yes. You can enable visual input only where it improves the visitor experience and keep other assistants text-only. That keeps upload behavior easier to review and avoids adding complexity to simple answer paths. The operational question is whether vision makes the workflow clearer once real conversations, real ownership, and real edge cases show up. That is the bar teams should use before they expand the rollout across more assistants, more channels, or more teams.
When should vision hand off to a person?
Hand off when the image is sensitive, ambiguous, outside approved source coverage, or tied to a decision your business wants a human to own. The upload and transcript can travel with the support or lead context. The operational question is whether vision makes the workflow clearer once real conversations, real ownership, and real edge cases show up. That is the bar teams should use before they expand the rollout across more assistants, more channels, or more teams.
Ready to get started?
Start your 7-day free trial. Review current trial terms.
7-day free trial