Follow me on LinkedIn - AI, GA4, BigQuery

New research from AVOXI and Metrigy says AI voice has gone mainstream in the contact centre.

55% of enterprises already use AI voice agents. 92% are using them or plan to within two years. 

But companies with customers in 37 countries, on average, have those agents live in only 17. 


Among the biggest firms, 71% already use voice AI, yet those operating in 50-plus countries have rolled it out in only about a quarter to a third of their markets.

Security and compliance are the main blockers (61%), followed by data sovereignty and privacy rules (51%). Language limits and gaps in phone or SIP infrastructure are also holding teams back.


Where AI voice is live, it is mostly helping people, not replacing them. 
81% use it for identity checks, 69% for tasks like payments and order status, and 68% to assist agents in real time. 

67% say it lets human agents handle the harder, more personal cases. Only 17% say they have replaced Tier 1 agents entirely.


The voice stack itself is still messy. 

60% of companies still use six or more voice providers, and 87% are consolidating or thinking about it. Most still run a mix of cloud and on-prem systems.

Outbound calling is another pressure point. AI is already in most outbound campaigns, but spam flags, compliance rules, and unknown numbers still stop people from answering. 


On security, companies worry about eavesdropping, weak VoIP APIs, and denial-of-service attacks. Robocalls and caller ID spoofing get treated as the highest operational risk.

More Details: https://www.businesswire.com/news/home/20260929187559/en/AI-Voice-Goes-Mainstream-Raising-the-Need-for-Global-Voice-Infrastructure-Orchestration-and-Security 


ElevenLabs launched Eleven v4 and Eleven v4 Turbo.

v4 is their most expressive text-to-speech model yet. It reads a script more like a voice actor than a typical AI voice. 

You can add direction straight into the text, like laughs, whispers, or a door slamming, and the model follows those cues more reliably than before.


It supports 90+ languages, multiple speakers in one clip, and built-in sound effects. 

Long scripts stay consistent from start to finish, so an audiobook can sound like one continuous take. 


If you regenerate a line, the voice doesn’t drift. 

v4 Turbo is the fast version for live conversations and voice agents. First speech comes back in about 150 milliseconds, so calls and chat agents feel more natural. 

You can use 17,500+ voices, clone a voice from a short clip, or describe a voice in a sentence and generate it. There’s also better control over how names and technical terms are pronounced.

Both models are available now, including through the API, and they work on the free plan.

More Details: https://elevenlabs.io/v4 


Scribe v2 Medical is now generally available.

ElevenLabs released Scribe v2 Medical (a medical version of Scribe v2), a speech-to-text model built for clinical audio. It is now available to everyone through the ElevenLabs Speech to Text API. 

On clinical audio, it cuts word error rate by about 35% compared with the standard Scribe v2 model.


Clinical recordings are hard to transcribe well. Drug names can sound almost identical, and notes often pack in dosages, units, and medical terms in quick succession. A mistake can end up in a patient record, so accuracy matters more than it does in everyday speech.


Scribe v2 Medical is HIPAA-eligible for enterprise customers with a Business Associate Agreement and Zero Retention Mode turned on. With that mode, audio and transcripts are deleted as soon as the request finishes.

It is billed at the same rate as Scribe v2.

Read the full announcement: https://elevenlabs.io/blog/scribe-v2-medical-generally-available 


Google added two new speech models to Gemini: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. 

In several public tests, they now sit ahead of ElevenLabs’ latest voices.

Gemini 3.8 Flash TTS is the creative one. You can design new character voices from scratch with a text prompt, then direct the performance line by line. Google points to games, audiobooks, podcasts, and other interactive media.

Gemini 3.8 Flash-Lite TTS is the cheaper, high-volume option. It is aimed at dubbing, lots of audio content, and voice agents, with control over tone, pacing, and expression.

What can you do with them?

>> Create custom voices by describing the role, accent, and personality. Google says this works across 100+ languages and dialects.

>> Use a library of 2,000+ ready-made voices, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.


>> Clone a voice from a 30-second sample, if you have the rights to use it. Google says this includes consent checks, SynthID watermarking, and C2PA credentials.

>> Save custom voices so they stay consistent across a project.

>> Voice remixing is coming soon, so you can tweak timbre, pitch, pace, and accent with a prompt. 


Where to try it?

>> Developers can use both models in the Gemini API and Google AI Studio.

>> Gemini 3.8 Flash TTS is also in Gemini Notebook.

>> Gemini 3.8 Flash-Lite TTS is in Google Vids.

>> Enterprise API access is coming soon.


Note(1): Voice cloning needs a matching verbal consent recording from the voice owner.

Note(2): All Gemini Audio output is watermarked with SynthID so it can be identified as AI-generated. Voice replication in AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland, or India.


Read the full announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/ 


ElevenLabs launched Studio 4.0 - an all-in-one AI editor for audio and video.

It’s built for video creators, podcasters, audiobook authors, and anyone who wants to turn a script or footage into a finished piece without bouncing between tools.


Here’s what it does?

>> Add natural voiceovers from 10,000+ voices. Change the text and the voice updates, no re-recording.

>> Generate custom background music that matches your scene.

>> Drop in sound effects from a simple description.

>> Clean noisy audio and isolate voices.


>> Auto-generate captions, including in other languages.

>> Edit audio and video on one timeline.

>> Use Studio Agent: describe what you want and it can draft a script, pick voices, add effects, and arrange clips. You can keep working with the agent or take over manually.


You can work in 30+ languages, share a project link for feedback, and export for personal or commercial use depending on the plan.

There’s a free tier to try it.In short: one place to write, voice, score, caption, and polish podcasts, videos, and audiobooks.

Try it: https://elevenlabs.io/studio 


ChatGPT Voice just got a lot more useful.

OpenAI has rolled out a big update. You can now talk to ChatGPT and actually get work done, not just chat.

Here’s what’s new:

>> Voice can use your connected apps. Ask it to check your calendar, look through email, or pull something from Slack without switching screens.

>> It’s now powered by the new GPT-6 models, Astra for the hardest jobs, plus Sol and Luna when you want something faster or cheaper.

>> Voice works inside ChatGPT Work on the web and on your phone. You can create documents, slide decks, websites, and spreadsheets just by talking.

The update is rolling out worldwide in the latest version of the ChatGPT app.


New research says a short voice recording could help screen for type 2 diabetes.

Researchers say 20 seconds of speech may help screen for type 2 diabetes.

An AI model trained on tens of thousands of voices flagged people with the condition most of the time, but it also produced a lot of false alarms.

It is not a diagnosis. The hope is that a phone recording could identify who should get a blood test.

More details: https://www.forbes.com/sites/fionariley/2026/09/29/ai-voice-test-could-screen-for-type-2-diabetes-research-finds/ 


AI Voice Clone Helps Drain €95 Million From Italian Bank.

Scammers used a fake WhatsApp account and an AI-cloned voice to get a senior executive at Italian private bank Fideuram to send about €95 million overseas. About €53 million has been recovered.

The rest was moved into crypto and is still missing.

More Details: https://www.yahoo.com/news/us/articles/ai-voice-clone-fake-whatsapp-145038517.html


Reception.ai is ElevenLabs ’ new AI receptionist for small businesses.

Reception by ElevenAgents answers the phone, books appointments, and handles after-hours calls. It is aimed at home services, professional services, automotive, wellness, and similar offices that live on inbound calls.

It is an ElevenLabs product, not a separate company. ElevenLabs launched it as “Reception by ElevenAgents.”

ElevenAgents is ElevenLabs’ platform for voice and chat agents; Reception is the ready-made receptionist built on that platform, using ElevenLabs voices.


What it does

  • Answers common questions on the spot: hours, pricing, availability.
  • Books appointments while the caller is still on the line, checks calendars, and writes the booking into the schedule.
  • Picks up nights, weekends, and holidays instead of sending people to voicemail.
  • Speaks in 70+ languages.
  • Offers an online booking page that stays in sync with phone and text, so nothing gets double-booked.
  • Routes calls around staff hours, or steps in as backup.
  • Can cover multiple locations with one system and the same brand voice.

The company says it runs on the same ElevenAgents platform already used for larger customer-support and scheduling work.

Case studies on the site include handling 88% of inbound patient calls and cutting call costs by 66%, first-line support for 35 million US customers, and resolving about 90% of home-care scheduling calls end to end.


Pricing (as listed):

  • Basic - $29/month: 75 minutes, 1 number, 1 call at a time.
  • Plus - $79/month: 275 minutes, 3 numbers, up to 3 calls at once.
  • Premium - $199/month: 1,000 minutes, 5 numbers, up to 10 calls at once, multi-location.

Extra minutes are billed per minute. Plans include Zapier, webhook, and MCP integrations.


Retell is introducing Credit-based billing.

Starting October 1st, the Retell platform will use credit-based billing.

Call minutes (voice, LLM, telephony) will come off a prepaid credit balance in real time.

When credits hit zero, new calls stop until you buy more or auto-recharge tops up the balance.


Phone numbers, extra knowledge bases, CPS upgrades, and extra concurrency will still go on your card at the end of the month.

That cycle runs calendar month to calendar month; mid-month add-ons are prorated.

Credits never expire and are not refundable. Billing is per workspace; each workspace needs its own payment method.

Turn on auto recharge if you do not want live calls cut off.

Details: https://docs.retellai.com/accounts/billing


Google launched two new Gemini voice models: 3.8 Live and 3.8 Live Extended Thinking.

They’re meant to make talking to Gemini feel more natural. Less like a voice assistant that pauses and more like a conversation.

3.8 Live is the everyday version. It’s cheaper to run at scale. You can speak to it, show it what’s on camera, switch languages mid-chat, and it can start a task in the background without stopping the conversation.

3.8 Live Extended Thinking is the heavier version. For harder jobs, it thinks out loud while it works (“let me check that…”)  instead of going quiet. It can walk through multi-step tasks as they happen.

Google says you can use them for things like live help with onboarding, playing chess from a video feed, turning a sketch into code, making bookings, or drafting a plan by voice.

They’re also showing up in Search Live and in Workspace tools like Docs, Gmail, and Keep.

The audio they generate is watermarked so it can be identified as AI.


Available now.

  • Developers: Gemini API and Google AI Studio.
  • Businesses: private preview in Gemini Enterprise.
  • Everyone: Search Live; the Extended Thinking model is also in Gemini Live and some Workspace features for paying subscribers.

More here: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/


OpenAI put GPT-Live-1 in the API.

GPT-Live-1 first arrived in ChatGPT Voice on 8 July 2026.

On 10 September 2026, OpenAI put the same system in the API for developers.

It is meant to feel less like a walkie-talkie and more like a normal conversation: you can interrupt, pause, or talk over it, and it can say “mm-hmm” while you think.


How it works.

Older voice systems waited for you to finish, then transcribed, thought, and spoke. GPT-Live-1 is full duplex: it hears while it talks.

It hands harder work (reasoning, search, tools, coding) to a separate backend model, so the voice doesn't freeze.

In ChatGPT, that backend started as GPT-5.5; in the API, you choose the thinking model (OpenAI names options such as GPT-6 Astra). The API version doesn't support image input.


In ChatGPT, two versions rolled out worldwide on iOS, Android, and the web:

  • GPT-Live-1 - default voice for Go, Plus, and Pro.
  • GPT-Live-1 mini - default for free users.

In the API:

  • Voice layer costs $0.05 per minute, billed per second.
  • Backend model and tools are billed separately.
  • 12 new voices at launch: Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder.
  • Audio and text in and out; no images or video.

OpenAI says GPT-Live-1 scored about 30 points higher than GPT-Realtime-2.1 on Full Duplex Bench (turn-taking, pauses, interruptions).

Language-learning app Speak reported almost 80% fewer interruptions than its old turn-based system.

Paired with GPT-6 Astra, OpenAI also reports it leading Tau3 voice-agent tasks.

Details: https://openai.com/index/introducing-gpt-live-1-in-the-api/


StepFun’s audio team released StepAudio 3 Gen: one model for many kinds of sound.

The StepFun-Audio Team’s new paper describes StepAudio 3 Gen, a single audio model that can do text-to-speech, design a new voice from a description, sing, make sound effects, make music, and mix those together.

Most recent “do-it-all” audio models generate sound with diffusion.

This one treats audio more like language: it chops sound into discrete tokens and predicts the next token. Speech, music, effects, and singing all share the same token system.


They built it on a text language model and trained it in stages so adding audio wouldn't wipe out the model’s language skills.

Instructions use a simple script format: who is speaking, what the scene should sound like, and what happens in order.


They say it is state of the art on TTS and voice design, and still strong on speech, vocals, effects, and music.

StepFun also has a public hub for the wider StepAudio 3 lineup (Realtime, Music, Gen, ASR, TTS): https://static.stepfun.com/blog/stepaudio3/

Samples: https://stepaudiollm.github.io/step-audio-3-gen/

Paper: https://arxiv.org/abs/2609.12945


Wispr Flow has launched Canto, a speech model built for messy, real-world dictation.

Wispr Flow is an AI voice-to-text app. You speak, and it types polished text into whatever app you are already using.

On 17 September 2026, Wispr’s Advanced Interfaces Lab introduced Canto, its latest real-time speech model.

Most speech models look strong on clean studio audio, but people actually dictate through laptop mics, earbuds, and headsets, often with traffic, music, other voices, or whispered speech in the background.

Canto is trained for that world.


Canto is the first model in a larger family.

Wispr is already training a successor more than ten times larger, aimed at harder audio, better multi-speaker recognition for its notetaker, smarter use of personal vocabulary, and more languages.

Longer term, they want one model that both transcribes speech and works out who is speaking.

Full post: https://wisprflow.ai/canto


Retell is cutting LLM prices and raising some European call rates.

Voice LLM rates on GPT, Claude, and Gemini are down now. Biggest voice drops: Gemini 3.6 Flash 65%, Gemini 3.5 Flash Lite 58%, GPT 5 nano 47%.

Chat and Fast-tier rates are also down for the same families.


From 1 October, outbound calls from Retell Twilio numbers to four countries go up (per connected minute):

Germany: $0.10 → $0.60

Spain: $0.10 → $0.20

France: $0.06 → $0.25

Italy: $0.06 → $0.30

Details: https://www.retellai.com/pricing


Retell is retiring the old built-in Cal.com tools.

check_availability_cal and book_appointment_cal are being replaced by the Cal.com integration.

One workspace connection replaces a copy of the API key on every agent.

The new integration also supports list, get, reschedule, and cancel and EU-hosted cal.eu accounts, which the old tools never did.


Dates:

  • 23 August 2026 - dashboard already stopped offering the old tools
  • 30 September 2026 - API stops creating or editing them. Last day to change a key, event type, or timezone
  • 31 October 2026 - Retell migrates existing tools to the integration

Migrated tools keep their function name, event type, and timezone.

Tools with an expired or revoked key are skipped and will fail on the next booking call.


What to do.

  1. Connect Cal.com once (use a key that never expires).
  2. Add Check Availability and Book Appointment to each agent, matching the old event type ID.
  3. Delete the built-in tools. Test before you ship.

Details: https://docs.retellai.com/deprecation-notice/2026/10-31_legacy_calcom_tools