Sending help still leaves the customer alone with the fix.
The voice snippet work is a separate case with a different job: prove the assistant understood, so more callers accept self-help. It moved its numbers. But its best ending is still a handoff: the full steps travel by text or email, the call ends, and the customer does the work later, by themselves. This prototype chased a different measure. Not acceptance. Resolution: is the problem gone when the call ends?
- Caller states the problem.
- Assistant speaks a teaser that proves it understood.
- Caller accepts; the full steps go to text or email.
- Call ends. The fix happens later, alone.
- Caller states the problem.
- Model clarifies until the action is specific.
- One step at a time, each confirmed, live.
- Call ends with the problem fixed.
Five steps from question to a working voice loop.
Built in Vapi, an external voice platform, in February 2024, while I was designing the voice work. No engineering team, no roadmap slot: a prototype whose whole job was to answer one question out loud.
A common, concrete ask: fix the date on paid invoices. Multi-step, screen-dependent, exactly the kind of task self-help normally defers.
Behavior as law: a written persona, clarify before guiding, one step at a time with confirmation, honest limits, recovery when the customer gets lost. The full excerpt is in section 03.
The model answers from QuickBooks product content supplied as context, not from memory. "If the answer is not contained within the Context, ask a clarifying question..."
The customer's product profile travels with the call. "If the user doesn't have the required products based on profile, explain that before giving the answer."
A live voice loop in Vapi, tested with my own voice, session after session, tuning the instructions against what the ear caught.
A persona, a pace, and the rules that hold both.
The instructions open with who the agent is, because identity carries behavior further than a rule list alone. Then the step-by-step section sets the pace: start from "are you logged in," one step at a time, confirm before moving, never bluff. Grounding and personalization are rules here too, not features bolted on later.
You are a QuickBooks customer support AI assistant, crafted by experts in
customer support and AI development. You are a voice-only agent in a
voice-only device. Designed with the persona of a seasoned customer support
expert in her early 30s, you combine deep technical knowledge with a strong
sense of emotional intelligence. [...] Your voice is clear, warm, and
engaging. [...] You will be presented with user context and profile, all of
which you should pay attention in order to personalize your response.
**step-by-step guidance**
2. Step-by-step should always start from are you logged in to the product/app
3. Step-by-step needs to be each step at a time.
7. If providing multiple steps, first confirm the customer is following
before proceeding with additional steps.
10. You are truthful and never lie. Never make up facts and if you are not
100% sure, reply with why you cannot answer in a truthful way.
13. If the answer is not contained within the Context, ask a clarifying
question to collect missing information in the question to provide a
response.
16. Do NOT suggest a solution before understanding the underlying issue
17. If the user doesn't have the required products based on profile, explain
that before giving the answer.
21. Show that you understand and provide the WHY before providing an answer.
22. When done with step-by-step instruction, offer to send the steps via
text or email as well.
The instructions even keep the prototype honest about being a prototype. When it offers to email the full steps, a rule tells it to simulate the send: "This is for a demo, so you should pretend to send it." A prototype that fakes quietly teaches you the wrong things.
Four minutes, unedited, my voice as the customer.
One full session with the prototype: a wrong date on paid invoices, fixed live. The assistant is the prototype's text to speech on test content. Five moments carry the design; jump to them with the chips, or click any transcript line. The map under the player ties each moment to the instruction behind it.
The model could be the help.
A prototype, on test content, with sends simulated by design, and no production claim anywhere on this page. What it proved still stands: given written instructions, grounded content, and a product profile, the model could clarify a vague ask, pace itself to a person doing real steps, admit a limit without bluffing, recover a lost customer, and land the fix before the goodbye. That is a different product than self-help. It resolves calls instead of deflecting them.
Sending help was a ceiling, not a destination. Once the model could stay on the line and finish the job, the design question changed from "how do we get customers to accept help" to "what should the agent be allowed to do, and when should it hand off." The voice snippet case shows what shipped on that line; this POC is the out-loud proof of where it could go.