You're standing in the kitchen, hands covered in dough, and you want to know how much flour is still missing. Old way: unlock the phone with your elbow, type with a knuckle, curse. Now you just ask out loud — and the AI answers while you keep kneading.
Talking instead of typing. That's what the AI companies are grinding on hardest right now. And over the past few weeks, it's gotten genuinely good.
What changed
On July 8, OpenAI released "GPT-Live," quietly replacing the old "Advanced Voice Mode." Google has had Gemini Live for a while now, which inherited the throne from the aging Google Assistant. Both want the same thing: a conversation that feels like a conversation — not commands barked at an answering machine.
The magic word is "full-duplex" (listening and speaking at the same time). Sounds technical, means something very ordinary: you can cut the AI off mid-sentence. "No, wait, I meant —" and it stops talking and listens. Like a human. The old voice mode couldn't do that — you had to wait for the monologue to finish.
Honestly, talking to an AI model on the phone isn't even new. Anyone self-hosting OpenWebUI (a graphical front-end for local and cloud models) could already call up practically any major model through its Conduit app — via OpenRouter, a single gateway to nearly every model in the world, even almost all of them at once. The catch: billing was per token, and the models weren't tuned for live replies. The listening and speaking — STT and TTS, speech-to-text and back — sat in the app or an upstream workflow, not in the model itself. Result: a noticeable wait before the answer came.
That's the real leap with GPT-Live and Gemini Live: the voice now lives inside the model, not bolted on around it. That's why it feels fluid instead of like a walkie-talkie on a delay. (If you'd rather go the self-hosted route without leaning on Big Tech, our OpenWebUI article shows how.)
What does it cost — and what's free?
Good news first: talking to AI is free.
On ChatGPT, every free account now gets "GPT-Live-1 mini" — the smaller model for chit-chat, dictation, and hands-free questions around the house. Its bigger sibling "GPT-Live-1" stays reserved for paying customers. For "how long do lentils take to cook?" the mini is plenty. When the question gets trickier, it quietly routes to a stronger model in the background — you won't notice.
On Google, Gemini Live is simply part of the Gemini app and usually pre-installed on Android. Free, deeply woven into the phone, and it can even look through the camera or read along on your screen.
In short: trying it costs nothing — except the few minutes it takes to get over talking to your phone like it's a person.
What you can do with it right away
The best of it comes out when your hands are busy. Cooking, crafting, wrenching on something — you can still ask. Practicing a language works surprisingly well: the AI chats with you in French and corrects you along the way, no blushing required. And anyone with kids knows the endless loop of "why" questions — a patient language model answers those without complaint.
Speaking of limits: live translation is the part where the marketing shine flakes off fastest. In early tests, GPT-Live's Hindi sounded like an American reading a menu for the first time. Fine for holiday small talk. For an important conversation: not yet.
The catch nobody likes to mention
For the AI to answer instantly, the microphone has to listen. Not secretly and around the clock — but the moment you start the mode, the channel is open, and everything you say goes to a server in some data center.
That's no reason to panic, but it is one to think. What you tell the AI while cooking is harmless. What you casually tell it about your boss, your health, or your tax return — because talking feels looser than typing — is a different league. The microphone lowers the barrier. Convenience and data thrift pull in opposite directions here. (If you want the details: what happens to my data?)
Two concrete things you can do: clear the voice history in settings regularly, and keep typing the sensitive stuff instead of saying it. No sacrifice — just a bit of switched-on brain.
Is it for you?
If you've only ever used AI by typing: yes, give it a go. It's the lowest-barrier way to bring AI into everyday life — no prompt craftsmanship, no blank text box, just talk.
If you always found voice assistants like Siri or Alexa a bit silly: this is a different class. The difference is roughly the one between a satnav barking "make a U-turn" and someone in the passenger seat who actually thinks along.
And if you stay skeptical: also fine. Nobody has to talk to their phone. But asking "explain quantum physics like I'm eight" out loud into the kitchen — and getting a patient answer while the pasta boils — that's got something to it.
Talking is the oldest user interface in the world. Maybe it's also the best.
