I turned voice on by default. That changed how I think about news
After a few weeks with OpenAI’s GPT-Live, I think voice is about to change how we access information, not just how we talk to machines.
I’ve had early access to OpenAI’s new voice model for a few weeks, and I ended up doing something I didn’t expect: I turned voice on by default.
I had mixed feelings about voice interfaces before. They were impressive in demos, but not natural or useful enough for me to start there. And it’s not that I wasn’t trying hard to like it: the first demo of Mizal was actually a voice assistant for news, almost two years ago.
This one changed that.
OpenAI launched GPT-Live today, and the first reactions captured the moment pretty well. “2026 is the year voice agents finally became good”, says this tweet. Olivia Moore at a16z called it “a big, big step up for consumer voice,” while pointing out that naturalness and reliability had not been good enough to foster heavy use before.
That feels right. The little frictions have started disappearing: GPT-Live can listen and speak at the same time. It handles interruptions, pauses and changes of pace more naturally. It is better with background speech, can switch languages mid-conversation and handles my imitation of “fo shure” beautifully. It can stay quiet while you think instead of treating every silence as the end of your turn.
It is not perfect, but it feels natural enough that you can have a real conversation. And once that happens, voice stops feeling like some bolted-in weird robotic additional feature and becomes the point of entry.
And this is not only an OpenAl product story. A few days ago, Hugging Face published a demo that points in the same direction from the open infrastructure side. Their stack is modular (speech recognition, fast inference, text-to-speech), so each layer can be swapped and inspected.
The stupid question test
All of that started to change how I access information and news.
When I read an article, watch a video or listen to a podcast, I follow the path created by the author. That is often exactly what I want. A good artifact has structure, judgment and an argument.
But sometimes I just want to get updated or oriented - you know, the good ol’ user needs model. What happened? Why does this matter? Wait, who is that person again?
These are not sophisticated questions. Some are the slightly stupid questions I would never dare to ask to another human being. But they are often the questions that help information click. It’s the “intimacy dividend” described by Shuwei Fang last year, on steroids.
With voice, you can ask one question, interrupt the answer, go sideways, come back and keep moving. The cost of a follow-up becomes almost zero.
You can also choose how much reasoning you want. For news, I usually select the highest level. It adds a little latency when the system needs to search or think, but in my experience the answers are more nuanced and better grounded.
That changes the experience from “give me a summary” to “let me ask my way into this story until it makes sense.”
Read it for me
This is where things get more consequential for news.
For the “update me” part of my information diet, the article will probably matter less. Not because an AI answer should replace original journalism - answers are grounded in it. But the artifact may no longer be the first interface I use.
A deeply reported investigation is still an artifact worth reading (and it can also benefit from an AI voice assistant, as Alessandro Alviani shared recently). A video can show evidence and emotion that text misses.
But if I want to understand why a central bank changed rates or why a World Cup matchup feels closer than the odds suggest, I may increasingly begin with a conversation with AI.
For that kind of usage, an article will be a piece of context among others. That raises all kinds of questions we started to scratch at with previous iterations of AI products: attribution, monetization, structured data vs. articles, and so on.
A new first gesture
For years, the voice debate has often been framed as a replacement question: will talking replace typing? Will voice replace the screen? Let’s frame it differently.
Voice may replace the first gesture.
Instead of opening an article, scanning a feed or carefully composing a search query, you begin by asking what you want to know. The conversation helps you orient yourself.
Interestingly enough, it also raises the profile of the screen - this surface of a few inches that has served us so well.
There are lots of bets being made on no-screen devices. Some better than others (looking at you, Snap Specs). But let’s not forget voice isn’t a perfect form factor though. Alone, it’s bad at discovery. You cannot easily skim a spoken answer, compare five numbers or jump between sources. Audio unfolds in time; a screen lets you inspect the whole thing at once.
This is where the integration of voice in ChatGPT is clever. When I asked questions about World Cup matches, voice let the conversation move naturally and the screen gave me the things audio is bad at presenting: visual cards and stats I can explore.
I can talk through the question, then glance at the structure. That combination feels much closer to how people actually try to understand information.
The screen is not a fallback for when voice fails. It is the evidence layer.
Oral on the surface, literate underneath
Allow me to be slightly philosophical for a moment, so we can think about what it means for how we will design information and entertainment products.
In a recent Atlantic conversation about Walter Ong’s work, Derek Thompson points to a limitation of written text: it is fundamentally unresponsive. You cannot ask an article to clarify itself, challenge one of its claims or explain what you are missing. It can only give you the words already on the page.
AI changes that. It turns the products of literate culture - articles, transcripts, documents and archives - into something we can question conversationally. Voice makes that shift literal.
But conversation alone is not enough for news. We still need what writing gave us: a durable record, evidence we can inspect and enough distance to think critically. The opportunity is to build products that are oral on the surface and literate underneath.
Two months ago, I wrote about three possible futures for voice AI. One of the strongest arguments against voice becoming the primary interface was simple: most of what we do on our phones involves browsing, glancing and choosing. Voice is terrible at that.
GPT-Live does not solve the problem by replacing the screen. It makes voice and the screen part of the same interaction. That may be the more important design shift.
The hard part for news is not speech
Voice could become part of the baseline against which other products are judged, including products that have nothing to do with OpenAI. The same way users started to expect to ask questions in natural language and get a useful answer immediately after ChatGPT launched.
That doesn’t mean we should have chatbots and AI voice everywhere. We need to figure out where the real value is for users. But it should push publishers and creators beyond the question of how to make our content available in voice, to explore what becomes possible when voice is finally good instead.
Imagine:
A live election, earnings or sports companion that talks through developments while updating charts, probabilities and sources on screen.
A podcast you can question, with answers linked back to the original clips.
A topic expert built around a newsroom’s reporting that can explain a story at different levels without flattening the uncertainty or disagreement.
Will the audience be interested in AI voice products that are not general assistants? The only way to get the answer is to try. There is at least something more defensible in the depth and the uniqueness of the experience, in addition to a deeper vertical integration.
It matters because it’s even more engaging than text-based chatbots, and I suspect that translates into higher engagement time and richer audience signals.
The challenge for news will also be to make reporting legible inside a conversation without losing where the information came from. If an assistant combines five articles, a podcast, a market signal and a public document into one smooth answer, the user still needs to understand:
Which source supports which claim?
Where do the sources disagree?
What is reported fact, analysis or prediction?
What should I open if I want to inspect the evidence myself?
The more natural the answer feels, the easier it is to forget that it is a synthesis produced by a system that can make mistakes. That makes provenance and source design part of the product. And you can’t have footnotes with voice!
Read more
What 13 voice AI insiders can’t agree on
Voice could change how we interact with machines forever. It’s the obvious UX layer for the agentic future: how we’ll dictate to agents, brief AI assistants, run our work without keyboards. Or it stays what it is today: a feature we use occasionally, mos…



