Feature

Balance answer quality behind the brand

Use owned content to answer visitor questions with less friction.

Start free trial

What this feature covers

  • Provider choice
  • Grounded answers
  • Cost control

Context

Why it matters

The practical reason to use it.

Multi-model support is useful when it improves a real visitor experience, not when it turns the product into a provider comparison page. InsertChat lets teams choose the right model profile behind a branded assistant without changing the website experience visitors see.

That gives operators room to balance speed, cost, and quality. A high-volume support flow can use a lighter model, while a research or escalation path can switch to a stronger one without changing the rest of the assistant setup.

Model choice is an operating decision that affects performance, budget, and trust. It should help the assistant answer website questions well, not distract visitors with technical choices they did not come to evaluate.

Teams also need a page that explains what stays constant while models change. The prompt, retrieval layer, tools, analytics, and handoff rules should remain stable so operators can compare model behavior on equal footing. That makes it much easier to answer practical questions like when a premium model is worth the spend, where a fast model is enough, and how multimodal requests should be routed without breaking the visitor experience.

Models usually gets prioritized when the current workflow is already creating manual review, unclear ownership, or brittle handoff between teams. The feature matters because it tightens the operating model around the assistant, not because it adds one more box to a feature matrix.

Teams need to understand how to launch the feature safely, measure whether it removes friction, and decide when the rollout is ready to expand. Before launch, record a baseline for completion time, escalation rate, manual corrections, and unresolved conversations. Compare those signals after the first bounded deployment, investigate regressions, and widen access only when the feature improves the downstream workflow without increasing review overhead. Those operating details make the capability concrete enough to evaluate before deployment.

Models needs a clear post-launch review. Operators should measure whether the feature reduces manual work, improves handoff quality, and stays predictable when real traffic and exceptions hit the workflow.

That review path is what keeps models from becoming another checkbox feature. Teams need enough detail to see which signals matter in production, where escalation still belongs, and how the rollout expands without losing control of quality.

How it works

How it works

A step-by-step look at the workflow.

  1. Step 1

    Choose a default model for the assistant based on the task and traffic profile.

  2. Step 2

    Switch to a different model when a conversation needs more speed, depth, or multimodal support.

  3. Step 3

    Keep the same knowledge and tools available so the assistant stays consistent across model changes.

  4. Step 4

    Review cost, latency, source use, and answer quality signals to refine which model should handle each question type.

  5. Step 5

    Turn those routing decisions into a repeatable rule so the team can improve quality without changing the website launch experience.

Coverage

Core job

The model layer is useful when teams can compare providers and tiers without redoing prompts, retrieval, branding, or handoff logic.

OpenAI models

Use OpenAI options for stronger reasoning and flexible routing when a visitor question needs more depth.

Anthropic models

Use Anthropic options for nuanced writing, long-context source review, and customer-facing responses where tone and reliability matter.

Google models

Use Google options when documents, images, and multimodal reasoning support the same grounded assistant.

Open and alternative models

Use additional providers when teams need cost flexibility, portability, or a different reasoning profile for a specific assistant route.

Coverage

Daily use

Routing matters when common, complex, visual, and long-context questions need different quality, speed, or cost choices.

Switch mid-conversation

Change models without losing chat history, retrieved context, or enabled handoff paths when the question needs a different response profile.

Cost control

Use lower-cost models for repetitive answer paths and reserve stronger models for questions where poor guidance would create follow-up work.

Assistant defaults

Set default models by assistant so support, product discovery, and content guidance each start from the right quality target.

BYOK support

Bring your own API keys when procurement, billing, or provider governance requires the model relationship to stay directly under your own vendor account.

Coverage

Control points

Teams get more value from models when rollout ownership, review, and downstream handoff stay visible after launch.

Start with one bounded workflow

Use Models on the narrowest workflow where the team can measure whether the feature reduces friction, improves clarity, and creates better cost control with model flexibility without adding extra review overhead. That bounded launch makes it much easier to see which inputs, rules, and team habits still need work before the capability spreads to more assistants or customer touchpoints.

Keep edge cases visible

Review the conversations, prompts, and system actions tied to models so operators can see where the rollout still depends on manual judgment or incomplete source coverage. Document those edge cases directly because operational trust disappears when a broad capability hides the hard parts of deployment.

Connect the surrounding systems

Models is stronger when it sits beside the knowledge, integrations, and routing rules that determine what happens after the first answer or action. Treat it as part of a connected system, not as a standalone toggle expected to improve every workflow on its own.

Expand after the first proof

Once the first deployment is stable, teams can extend models into more surfaces and assistants without rebuilding the same control model from scratch every time. That is what lets a feature graduate from a nice idea into a repeatable operating pattern the whole organization can use with confidence.

Outcomes

What you get

The changes teams should notice first.

  • Better cost control with model flexibility
  • Higher quality for complex conversations
  • Faster responses with optimized model selection
  • More provider flexibility without changing the visitor experience

Proof you can check

The facts do the selling

Plan facts, platform capabilities, and worked examples — every claim here is checkable, not a pitch.

White-label included — never a paid add-on. Copyright removal from $98/mo. Full white-label — custom domain, branded portal, your-domain emails — from $198/mo.

The white-label wedge

Platform fact

Training runs on your sitemap, PDFs, docs, and YouTube transcripts. Answers cite the source pages they came from.

Trained on your content

Platform fact

Five clients at $300/mo on a $198/mo Agency plan is $1,300+ of monthly margin before usage.

A 5-client agency on one flat plan

Worked example

Questions and answers

Common questions

Practical answers about balance answer quality behind the brand.

Can I switch models without rebuilding the assistant?

Yes. The assistant configuration, approved sources, brand settings, and enabled handoff paths stay in place while the serving model changes. That lets teams compare providers without rebuilding the website experience. The operational question is whether models makes the workflow clearer once real conversations, real ownership, and real edge cases show up. That is the bar teams should use before they expand the rollout across more assistants, more channels, or more teams.

Why use multiple models instead of one?

Different visitor questions need different trade-offs. Multi-model support lets you use lighter models for common answers and stronger models for complex, long-context, or multimodal questions while keeping source grounding and brand behavior consistent. The operational question is whether models makes the workflow clearer once real conversations, real ownership, and real edge cases show up. That is the bar teams should use before they expand the rollout across more assistants, more channels, or more teams.

Does multi-model support help with cost control?

Yes. Teams can route traffic to the least expensive model that still meets the quality target, then escalate only the conversations that justify deeper reasoning or richer multimodal capability. That keeps model cost aligned with the business value of the request instead of treating every chat like the most expensive possible workload. The operational question is whether models makes the workflow clearer once real conversations, real ownership, and real edge cases show up. That is the bar teams should use before they expand the rollout across more assistants, more channels, or more teams.

Related resources

Ready to get started?

Start your 7-day free trial. Review current trial terms.

Start free trial