Cutting first-response time from 14 hours to 40 seconds
Three support reps were drowning in 2,400 tickets a month. We deployed a retrieval-grounded support agent that now resolves two thirds of them outright.
- Client
- A 40-person B2B SaaS company
- Industry
- Software
- Timeline
- 5 weeks to production
Results
68%
Tickets fully resolved by AI
Measured over 90 days post-launch across 7,100 tickets, excluding spam.
40s
Median first response
Down from 14h 12m, including overnight and weekend volume.
+22
CSAT points
From 71 to 93 on a 100-point scale over the same window.
The problem
What we found before anything was built.
- 01
Ticket volume grew 3× in a year while the support team stayed at three people. Backlog became permanent.
- 02
Roughly 70% of tickets were the same twelve questions — password resets, invoice copies, seat changes, integration setup.
- 03
Nothing was answered outside 9–5 Pacific, so international customers routinely waited a full business day.
- 04
Reps spent their mornings triaging rather than solving, and the genuinely hard tickets aged the longest.
The solution
What we built, and why it was built that way.
- 01
Indexed 480 help-centre articles, 18 months of resolved tickets and the internal product wiki into a vector store, chunked semantically and reranked at query time.
- 02
Built a support agent that answers only from retrieved sources and cites them. If retrieval confidence falls below threshold, it does not guess — it escalates.
- 03
Gave the agent four real tools: look up a subscription, resend an invoice, adjust seat count, and generate a scoped integration setup guide.
- 04
Wired confidence-based escalation into Intercom with a written summary, attempted steps and the customer’s plan and history attached.
- 05
Shipped a weekly evaluation run against 200 golden tickets so regressions surface before customers find them.
Architecture
How a single request moves through the system
- 01Ticket arrives
- 02Classify intent
- 03Retrieve + rerank
- 04Draft grounded answer
- 05Confidence gate
- 06Resolve or escalate
Every step is idempotent and retried with exponential backoff. Failures land in a dead-letter queue with a Slack alert rather than disappearing, and the confidence gate routes anything uncertain to a human with full context attached.
“We stopped hiring for the queue. My team spends their day on the problems that actually need a human, and our international customers finally get answers before they go to bed.”
Find out what your team could stop doing by next month.
Bring one process that frustrates you. We’ll tell you whether it can be automated, roughly what it costs, and what it would give back.
We map where the hours go
A quick walk through your day-to-day to find the processes eating the most capacity.
We rank by return
You get an honest read on what’s worth automating first — and what isn’t worth it at all.
You leave with a plan
A written recommendation and a rough number, whether or not you work with us.