Consumer & RetailGenerative AI · RAG

Support that cites its sources: a grounded RAG assistant for retail

A retrieval-grounded assistant that answers from the brand's own help center, policies, and live order data — deflecting routine volume and handing nuance to humans with context attached.

EngagementFixed-scope build
Timeline to v14–8 weeks
PatternRepresentative engagement
↑ deflection
on repeat questions
↓ first-reply
time
0
hallucinated policies

The problem

A consumer retailer's support queue was 60% repetition — where's my order, how do returns work, does this ship internationally — but every generic chatbot they'd tried made things worse by answering confidently and wrongly. Meanwhile the nuanced 40% waited behind the repetitive 60%, and agents burned out answering the same question for the thousandth time.

The hypothesis

The failure of the previous bots wasn't the model — it was groundlessness. Our hypothesis: an assistant restricted to retrieved, current, brand-owned sources — help center, policy docs, and the customer's live order data — with citations on every answer and a fast lane to humans, could deflect the repetitive volume without the trust-destroying errors.

The build

  • Grounding layer — help center and policy content indexed for retrieval, re-synced on publish, so the assistant can never cite a stale return policy.
  • Order-aware answers — authenticated customers get answers computed from their actual order state, not generalities.
  • Refusal + handoff design — if retrieval confidence is low or the topic is sensitive (payments disputes, complaints), the assistant says so and hands off to a human with the full conversation and retrieved context attached — no customer repeats themselves.

Design choice that mattered: we made "I'm not sure — let me get you a person" a first-class answer, tracked as a success metric rather than a failure. That single reframe is why customer satisfaction rose instead of dipping.

Rollout

Two weeks in shadow, answering silently alongside human agents while we measured answer agreement. Then a traffic ramp — 10%, 30%, full — with weekly review of every low-confidence handoff to expand the grounded coverage deliberately.

Results

Routine volume deflected without the accuracy incidents that killed prior attempts; first-reply times fell for the nuanced tickets that reached humans; and the citation-everything design meant the support team could audit any answer in one click.

What we'd tell you

  • Grounding beats model choice. A modest model with great retrieval outperforms a frontier model guessing.
  • Design the handoff as carefully as the answer — it's where trust is won.
  • Content ops is now part of support ops: the help center is the assistant's brain, so keep it current.
← All success stories
Keep reading

More from the field.

Banking & Financial

Consumer lending, decided in hours

60%
faster time-to-decision
files per underwriter
Read the story
Banking & Financial

Early-warning credit risk

60 days
earlier risk signal
↓ roll rates
into later buckets
Read the story
Banking & Financial

The 6 a.m. advisor brief

advisor capacity for client time
6 a.m.
brief ready daily
Read the story
Get started

Bring us a hypothesis. Leave with a system.

Tell us what's eating your team's time. We'll give you an honest read on whether AI is the right tool — and if it is, a scoped v1 with a timeline and cost.