skip to content
ai · TSMC ADDS $100 BILLION TO ARIZONA CHIP BET, TOTAL HITS $265 BILLIONai · META WILL ALERT PARENTS IF TEENS DISCUSS SUICIDE WITH META AIai · ROBLOX'S "BUILD" LETS ANYONE MAKE A GAME FROM THEIR PHONE WITH AIbusiness-of-tech · APPLE RECLAIMS WORLD MOST VALUABLE COMPANY TITLE, NVIDIA BOTTLES ITconsumer-tech · GOOGLE ADDS YOUTUBE MUSIC, INSTACART & CANVA TO AI MODE SEARCHai · ZOOX REALLS ENTURE ROBOTAXI FLEET OVER SMOKE DETECTION FAILUREai · TSMC ADDS $100 BILLION TO ARIZONA CHIP BET, TOTAL HITS $265 BILLIONai · META WILL ALERT PARENTS IF TEENS DISCUSS SUICIDE WITH META AIai · ROBLOX'S "BUILD" LETS ANYONE MAKE A GAME FROM THEIR PHONE WITH AIbusiness-of-tech · APPLE RECLAIMS WORLD MOST VALUABLE COMPANY TITLE, NVIDIA BOTTLES ITconsumer-tech · GOOGLE ADDS YOUTUBE MUSIC, INSTACART & CANVA TO AI MODE SEARCHai · ZOOX REALLS ENTURE ROBOTAXI FLEET OVER SMOKE DETECTION FAILUREai · TSMC ADDS $100 BILLION TO ARIZONA CHIP BET, TOTAL HITS $265 BILLIONai · META WILL ALERT PARENTS IF TEENS DISCUSS SUICIDE WITH META AIai · ROBLOX'S "BUILD" LETS ANYONE MAKE A GAME FROM THEIR PHONE WITH AIbusiness-of-tech · APPLE RECLAIMS WORLD MOST VALUABLE COMPANY TITLE, NVIDIA BOTTLES ITconsumer-tech · GOOGLE ADDS YOUTUBE MUSIC, INSTACART & CANVA TO AI MODE SEARCHai · ZOOX REALLS ENTURE ROBOTAXI FLEET OVER SMOKE DETECTION FAILURE
BAD/GATEWAY*

OPENAI'S CHATGPT LIVE WANTS YOU TO TALK OVER IT, LITERALLY

The new voice model can listen and speak at once, peppering your sentences with "mmhms" and "sures" instead of dead air.

by editor6 min readcomments soon

openai's gpt-live wants you to talk over it — literally
· Image credit: OpenAI

OpenAI is replacing ChatGPT's voice mode with a model that can listen and talk at the same time. The new system, called GPT-Live, makes decisions about whether to speak, keep quiet, interrupt, or call a tool dozens of times per second. It is designed to end the stilted back-and-forth that has made voice chatbots feel like phone trees with anxiety.

The biggest change is visible in the smallest sounds. GPT-Live will drop acknowledgement phrases — "mmhm," "sure" — while the user is still speaking. Those micro-responses are what humans do naturally to signal attention. Their absence in earlier voice AIs created the signature dead space where users pause, expecting a response, and the AI remains silent until the utterance is complete. That beat of silence is the thing GPT-Live is designed to kill.

The model replaces the existing ChatGPT voice experience entirely. OpenAI describes it as a new generation of voice model built for continuous interaction. That means the AI is not waiting for its turn. It is listening, processing, and making a call at sub-second latency on whether to chime in, let you finish, or take the floor. The architecture is fundamentally different from the previous pipeline, which recorded a full utterance, transcribed it, generated a response, then played it back. GPT-Live collapses those stages into one flow.

THE PROBLEM WITH SILENCE

Every voice UI that relies on push-to-talk or strict turn-taking introduces a cognitive tax. The user learns to speak in discrete, self-contained sentences. That is not how humans talk. People pause, restart, trail off, and rely on the listener's backchannel cues to know whether they are being followed. A listener who gives nothing back forces the speaker to fill the silence, which is what happens with most voice assistants today. You say something, wait, and the lack of response pressures you to speak again or repeat yourself.

GPT-Live's approach is closer to what a human listener does. Acknowledgement tokens do not carry semantic weight. They are purely phatic signalling; the channel is open. But that signal changes the exper. The model's ability to decide when to use those signals, and when to stay silent or jump in with a real response, is learned from the same conversational training data that powers the underlying language model. It is a behavioural change with outsized effect on perceived naturalness.

LISTEN AND SPEAK AT THE SAME TIME

Full-duplex conversation — both parties talking simultaneously — is the other half of the trick. GPT-Live can process incoming audio while generating outgoing speech. That is not simply faster audio processing; it requires the model to handle overlapping streams without garbling either one. When a user interrupts the AI mid-sentence, the model hears the interruption and can decide to stop talking, yield the floor, or continue after a brief overlap. The earlier voice mode would have to finish its sentence and then wait for the user's full input, creating the stilted rhythm of a bad radio interview.

The internal architecture runs multiple decisions per second: should I speak now, continue listening, pause, interrupt, or invoke a tool? Those are not sequential passes. They are parallel evaluations happening inside the model as audio flows. The result is a system that behaves less like a chatbot with a voice layer bolted on and more like a conversational partner.

WHAT CHANGES FOR USERS

The most obvious effect is that conversations no longer feel like interrogations. A user who says will hear an "mmhm" midway through, letting them know the AI is following. The previous version would sit silent, parsing the entire utterance before responding, forcing the user into unnatural clarity or repetition.

There is also a subtle psychological effect: acknowledgement phrases make the user feel heard, which reduces the tendency to over-explain. Users who get real-time confirmation that the AI is tracking their meaning tend to give shorter, more confident inputs. That shortens the overall interaction and reduces the cognitive load on both parties.

The model can also interrupt appropriately. If the user says something that immediately triggers a clear tool ca — — GPT-Live can begin the action without waiting for the user to finish the sentence. In the old system, the user had to stop talking, wait for the AI to finish transcribing, and then see the timer start. Now the timer can begin before the user has finished the instruction.

THE UNCANNY VALLEY OF BEING HEARD

There is a risk here.

A voice UI that is too conversational can cross into the uncanny valley of false presence. If the AI says "mmhm" at exactly the wrong moment — or too frequently — it can feel patronising. The model's decision engine has to calibrate the frequency and timing of those backchannel cues to match the user's own conversational rhythm. Humans do this unconsciously. A machine doing it consciously might always feel a beat off, even if the latency is imperceptible.

OpenAI is likely aware of this. The model was trained on human conversations, so its backchannel timing should approximate human norms. But users will have different expectations. Some will want the AI to stay quiet until they finish. Others will want constant reassurance that the line is open. GPT-Live needs to handle that range, or at least default to a broadly acceptable middle ground. The early reviews will focus almost entirely on whether the "mmhm" feels right.

BUT WILL IT FEEL MORE HUMAN? SHOULD IT?

GPT-Live signals the direction every voice AI is heading: full-duplex, real-time, phatic. The old model of recording and responding is dying. The next generation of voice interfaces will listen while they speak, nod while they think, and interrupt when it makes sense. That is a much harder engineering problem than faster speech-to-text, and OpenAI is the first major player to ship it at scale.

The immediate effect for ChatGPT voice users is a product that feels less like a tool and more like a person on the other end of the line. Whether that is comforting or creepy depends on where the uncanny valley settles. But the era of the awkward pause is over.


what did you make of it?

share

more from ai