The phone rings at the BDC desk with a call already in progress. An AI voice agent has spent ninety seconds with this customer, confirmed who she is, established that her appointment was moved without anyone telling her, detected that she is now annoyed, and transferred her to a person.
The BDC agent has a transcript on screen and about four seconds to read it.
Why this conversation is hard
This conversation did not exist three years ago, and there is no training for it anywhere.
AI phone agents are now installed across thousands of rooftops, answering after hours and handling routine inbound. Published dealer case studies put containment rates around 62 percent. That is a real operational win, and it has a consequence almost nobody has priced in.
The easy calls were automated. What reaches a person is the residue: the escalation, the complaint, the complex request, and specifically the customer the AI transferred *because* it detected frustration it could not defuse.
So the average difficulty of every human-handled call went up, while onboarding stayed a script and a day of shadowing.
Three distinct skills are now required and none of them are taught. Absorbing context fast from a transcript instead of asking the customer to start over. Acknowledging what already happened without disowning the machine that did it. And handling someone whose irritation has an extra layer, because they have already explained themselves once and are braced to do it again.
The single worst thing an agent can say here is "can you tell me what this is about?" The customer already did.
There is a related objection worth meeting head on, because practitioners raise it constantly. Ask r/sales whether AI roleplay is worth it and the replies run skeptical: it sounds robotic, it isn't realistic, you need at-bats with real people. One commenter warned against buying anything in this category before piloting it with your own inputs and data, on the grounds that a canned demo is easily fudged.
That is fair, and the last part is the useful part. A prebuilt scenario library is a canned demo at scale. A scenario built from a call your own store took last Tuesday is not.
The most telling comment in that thread came from a moderator working in insurance, who said they would love something that let them pre-load their own common objections and score responses against them, but doubted it could judge tonality, confidence and accuracy. They were describing the product they wanted and assuming it did not exist. The thread is from 2023, which is roughly when that was true.
How the agent is built
The scenario starts mid-conversation, which is the configuration that matters. The persona arrives with history: a prior interaction already occurred, facts were already established, and she knows the store knows them.
Behavior is set so she reacts sharply to any request to repeat information she has already given. She is not hostile at the start. She becomes hostile if handled as a fresh call.
Proactive sharing is set low deliberately. She will not re-offer her account details, her appointment time, or the reason she called. The agent has to work from what is on screen.
Objections and priorities carries the real ones: "I already told the robot this," "why am I starting over," "can I just speak to whoever actually makes decisions."
Because the roleplay agent can be built from your own recordings, the strongest version of this scenario uses real transferred calls from your own AI vendor. The transcript the trainee sees is a real transcript, in the format their system actually produces, with the same gaps and the same imperfect summary.
That is the difference between practising the situation and practising *your* situation, and it is not a difference a prebuilt scenario library can offer.
What gets scored
Whether the agent opened by demonstrating they had read the context rather than asking for it. Whether the acknowledgment was specific, naming what happened, rather than a generic apology. Whether they took ownership of the store's error without blaming the automated system, which reads as an excuse to a customer who does not care which part of the dealership failed.
Then the resolution itself: whether the fix offered was concrete and time-bound, and whether the customer's original goal was actually met rather than deflected.
The workflow half scores what happened in the CRM and scheduler while the conversation ran: whether the record was located from the transferred context, whether the appointment was actually rebooked, whether the interaction was dispositioned correctly so the next person to touch this customer inherits the truth.
That last step is the one that gets skipped under pressure, and it is the one that causes the next bad call.
What the run shows
A good run is short. The agent opens with "I can see your appointment got moved to Thursday and nobody called you, and I'm sorry about that, let's fix it." The customer's temperature drops immediately, because the thing she was braced to fight about has already been conceded.
The instructive bad run is not rudeness. It is the agent who is warm, apologetic, patient, and asks her to explain the situation from the beginning. She explains it. She is now much angrier than she was when the AI transferred her, and the call runs four minutes longer.
On a conversation-only tool that call scores reasonably. The tone was good throughout. The failure was structural, and it happened in the first ten seconds.
What a manager does with it
A BDC director gets a number for something that was previously invisible. Transferred calls can be scored as their own category and compared against calls that started with a person, which shows whether the handoff itself is where satisfaction is being lost.
That is a useful conversation to have with the AI vendor, too. If a specific transfer type consistently arrives with insufficient context, that is a configuration problem on their side, and it is now evidenced rather than anecdotal.
For onboarding, this scenario belongs in the first week rather than the fourth. A new BDC agent will take a transferred call on day one whether or not anyone has prepared them for it.
And because the same rubric runs on live calls, a director can confirm that the handoff opening practised in training is the opening being used on the floor. That is the whole point. Practice proves capability. Only the live call proves behavior.
Bring five real transferred calls from your AI phone vendor. We will score them and build the roleplay out of the worst one.
Frequently Asked Questions
The customer has already explained herself to a machine that decided it could not help. The facts are established, she knows the store has them, and the worst opening available is "can you tell me what this is about?" The average difficulty of every human-handled call went up when the routine ones were automated.
Ideally from your own AI vendor's real transferred calls, which is the strongest version of this scenario. The trainee then reads a transcript in the exact format their system produces rather than a clean one written for training.
The first week. A new BDC agent will take a transferred call on day one whether or not anyone prepared them for it.
Yes, and that is the point for a BDC director. Scoring them as their own category and comparing against calls that started with a person shows whether the handoff itself is where satisfaction is being lost, which is also a useful and evidenced conversation to have with the AI vendor.
It raises it. Containment removes the easy calls and leaves the hard ones, so the remaining conversations got harder while onboarding stayed a script and a day of shadowing.
Table of Contents
Talk to Sales
Have questions about training and enablement for your sales, CS, support, or leadership team? Let's talk.
Talk to Sales






