Podcast

Voice Reimagined: The Rise of Voice AI

In this 25-minute episode of Twilio Talks On Air, Twilio’s Peter Bell sits down with AI expert Prof. Kate Devlin and AI engineer Matt Cortland to explore why speech is replacing screens, and what it takes to build Voice AI that respects human psychology and eliminates customer friction.

Twilio Talks On Air Series

 

Voice is having a renaissance

 

For years, technology forced humans to adapt to screens, clunky drop-down menus, and frustrating automated journeys. Today, consumer expectations have officially reached a turning point: 41% of UK adults are far more comfortable using voice features today, and 22% of all consumers—rising to 35% of Gen Z—now talk more than they type.

In this 25-minute episode of Twilio Talks On Air, Twilio’s Peter Bell sits down with AI expert Prof. Kate Devlin and AI engineer Matt Cortland to explore why speech is replacing screens, and what it takes to build Voice AI that respects human psychology and eliminates customer friction.

 

Key takeaways

 

 

1. The Renaissance of Voice & The Behavioral Shift

[02:55] Peter Bell: Welcome to Twilio Talks. My name is Peter Bell, and I head up marketing for Twilio in Europe. Today, I'm joined by Kate and Matt, who I'll introduce in a moment. But first, I want to look back a few years. "It's good to talk"—that was the killer tagline that Bob Hoskins delivered as part of a very long-running 1990s campaign for British Telecom. And you know what? He may have been onto something, because after 20-odd years of "don't call me, text me," great use of our thumbs, typing, and swiping, it turns out that we actually do like to talk. And there's a real renaissance happening.

Twilio recently conducted some UK research showing that 22% of all adults would rather talk than message. If you look at Gen Z, that rises to 35%. So we're here today to explore the renaissance of voice.

Joining me today are Kate Devlin, Professor of AI and Society at King's College London, and Matt Cortland, an entrepreneur, AI engineer, and creator of the Guinndex, where he had voice agents call pubs across Britain to find out the cheapest pint of Guinness. Kate, beginning with you: our research shows 41% of us are inherently more comfortable talking to our computers and phones than a few years ago. From a human-computer interaction (HCI) and psychology perspective, what is happening with technology that is driving this renaissance?

[05:05] Kate Devlin: There are a couple of things. First, voice is our most natural way of communicating. It has very low cognitive load because we don't have to think about what we're doing; it's expressive, clear, and effective. Second, technology can cope with voice a lot better than it used to. Having had years of voice assistants, we've become habituated to it. It is the perfect moment bringing all of those things together.

[05:36] Peter Bell: A lot of us think of personal voice notes, but what is happening in the workplace that is driving this upsurge in interest?

[05:55] Kate Devlin: With large language models handling conversational interfaces so well, you can do a lot more now without typing. You can dictate, and the software picks things up quickly and produces correct content. Previously, you had to manually correct or steer things when they went wrong. The fluidity of the technology is fundamentally changing how people interact.

[06:26] Peter Bell: When designing these systems, where should teams begin to avoid recreating the terrible software interface experiences of the past?

[06:45] Kate Devlin: Because voice is so natural, we must ensure commands are interpreted correctly and easily. Systems need to handle accents and support accessibility for people who speak differently. We also must prioritize privacy and security, which is becoming increasingly critical with threats like voice fraud.

2. Real-World Voice AI in Action: Building the "Guinndex"

[07:20] Peter Bell: Matt, bringing you in—you build real-world solutions. Could you tell us about the Guinndex and how you approach designing voice interfaces?

[07:44] Matt Cortland: The Guinndex started from trying to determine the cost of a pint of Guinness in Ireland and the UK today. Having previously owned pubs and restaurants, I knew scraping websites was unrealistic—the only way to get the data was to visit in person or call.

I designed an AI voice agent using Twilio, ElevenLabs, and Claude to call pubs and plainly ask what a pint costs, then mapped that onto a website. When designing voice interfaces, I put myself in the shoes of the person answering the phone: What would they want, say, and think? I train the agents recursively by having them call me so I can provide real-time feedback.

[09:28] Peter Bell: How far have we come from traditional IVR call centers where you press one for sales, two for service, and zero for a human?

[09:47] Matt Cortland: Voice AI technology has advanced very quickly into a medium we can use meaningfully. We're not all the way there yet, but it's getting very close.

3. Psychology, Digital Trust & The 200ms Latency Rule

[10:02] Peter Bell: Kate, from a social and philosophical viewpoint, how important is it that we declare to callers that they are speaking with a synthetic voice?

[10:16] Kate Devlin: It's critical because people do not like being deceived by machines. People are comfortable interacting with machines as long as they know upfront. Studies show that finding out belatedly that an agent is a bot reduces trust, increases disengagement, and lowers sales completions. Declaring it upfront sets expectations, making users far more tolerant of minor flaws or uncanny effects.

[11:00] Peter Bell: What differentiates a good experience from a bad one in the human brain?

[11:21] Kate Devlin: A good experience has rapid, synchronous transitions. In human conversation, there is roughly a 200-millisecond gap between turns as we anticipate the end of a sentence and respond. With voice agents, longer delays feel unnatural and uncanny. You have to close that timing gap as much as possible.

[11:53] Peter Bell: Is it purely that latency, or are there other subtleties?

[11:59] Kate Devlin: Voice conveys enormous amounts of information—warmth, depth, and tone. Getting that right involves choosing pleasing accents and reassuring phrasing.

[12:28] Peter Bell: Matt, how do you take on that challenge technically to build an experience that delights rather than frustrates?

[12:37] Matt Cortland: It requires constant testing, listening, and adjusting. It comes down to fine-tuning prompts, context, phrasing, accents, and trying to break the system before users do.

[13:10] Peter Bell: How did you manage those millisecond delays using Twilio and ElevenLabs?

[13:26] Matt Cortland: It's tricky because you must balance giving someone time to pause versus jumping in appropriately. It involves extensive experimentation with underlying LLMs (like Claude or Gemini), prompt design, and speech speed adjustments.

[14:09] Peter Bell: How do you handle context and memory so callers don't have to restart the conversation if a call drops?

[14:31] Matt Cortland: For example, in a voice agent I'm building for Turing Fest, the caller fills out an initial survey. When the system connects, it recognizes their number, confirms their identity, and pulls in their prior answers so they never have to repeat themselves. In a new field like this, you have to look at the experience from a 360-degree perspective and constantly test.

4. Bridging Theory & Practice: Advice for Redesigning CX

[15:18] Peter Bell: Kate, what empirical academic research can builders leverage during testing?

[15:32] Kate Devlin: Empirical studies on accent perception, trustworthiness, and word choice, as well as sentiment testing to see how people feel about specific interactions. Doing upfront research is often far more effective than relying solely on post-call surveys that fatigue users.

[16:24] Peter Bell: Matt, when a client comes to you with requirements, how do you move from the base blocks to execution?

[16:42] Matt Cortland: You have to understand the core objective, but avoid having "too many cooks in the kitchen". Recently, I recommended throwing out a client's initial approach to rework the phrasing and cadence, which increased their success rate from 3% to 17%.

[17:24] Peter Bell: Are there specific sectors or use cases where this works best?

[17:33] Matt Cortland: Outbound calls work best with businesses accustomed to phone traffic, like hospitality and restaurants. However, I'm a huge proponent of inbound calls where the user initiates contact, retains agency, and knows they are engaging with an AI.

[18:32] Kate Devlin: Agency and control are essential for user comfort.

[18:39] Matt Cortland: Transparency is non-negotiable. You must disclose when someone is talking to AI.

[18:48] Peter Bell: Otherwise, it crosses into the "creepy line".

[18:50] Kate Devlin: Exactly—the uncanny valley. If something sounds almost human but slightly off, people get creeped out.

[19:28] Peter Bell: If leaders are redesigning an existing IVR system, what is the single most important rule from an HCI standpoint?

[19:47] Kate Devlin: Building trust through transparency so the user knows what to expect.

[20:03] Peter Bell: And Matt, what is your top piece of advice?

[20:14] Matt Cortland: Do your research on changing regulations. Laws around voice AI are evolving rapidly across different territories, and you should always disclose AI use.

5. Bonus Quickfire Round: Pop Culture, Wishful Automation & Dream Voices

[20:41] Peter Bell: To wrap up, we have a quickfire round inspired by the classic British radio game—no hesitation, repetition, or deviation! First: what is your most memorable phone call from cinema?

[21:32] Kate Devlin: Liam Neeson in Taken saying, "I will find you and I will kill you," which sounds even more menacing in a Northern Irish accent.

[21:47] Matt Cortland: The final season of Stranger Things, when the girl escapes by hitting someone over the head with the telephone handset.

[22:01] Peter Bell: Next: wishful automation. If you could delegate one annoying admin task in your personal life to an AI voice agent, what would it be?

[22:20] Matt Cortland: An agent that argues with customer support on my behalf to win bill disputes. I recently set up Claude to negotiate refunds, and it has won back thousands of pounds.

[22:51] Kate Devlin: I want Matt's agent! Failing that, an agent that calls the mechanic and mimics the broken car noises.

[23:16] Peter Bell: Mine is the "electric monk" from Douglas Adams, who believes in things so you don't have to. Finally: if you could pick any celebrity or fictional voice to power your AI assistant, who would it be?

[23:44] Kate Devlin: Malcolm Tucker from The Thick of It—a very sweary Scotsman.

[23:52] Peter Bell: That's two of us!

[23:55] Matt Cortland: Moaning Myrtle from Harry Potter.

[24:00] Peter Bell: Thank you both for joining. For everyone listening, you can find out more about Twilio's Voice AI products and download our full research on UK voice attitudes at twilio.com. Thank you!