In a blind test, readers preferred an AI-written data story over the human journalist’s version, 3 to 1.
That’s one of the insights from Data2Story, a new paper out of Oxford and Stanford. It’s the kind of number that gets screenshotted and turned into a “journalism is over” thread.
But the closer I looked and the more I ran the thing myself, the more I think it points the opposite way.
As you know, I’m especially interested in AI agents teams and how we can collaborate with them. At Mizal, we ran a similar experiment with a newsroom in a box, aka team of agents that collaborates with journalists. These days, we’re running a World Cup experiment with a team of agents that choose what they want to read and watch, and decide each day which stories and angles to publish (example below).
At the end of the day, I think these experiments tell us more about where the human work actually lives.
7 specialized agents
The idea behind Data2Story is seven specialized agents that hand off to each other, like a small newsroom: a Detective that gathers context, an Analyst that runs the statistics in real code, an Editor that decides the narrative, a Designer, a Programmer that builds the page, and - the clever part - an Inspector that binds every published claim back to the exact source behind it.
I pointed it at WHO life-expectancy data. Minutes later I had a finished, interactive, fully-sourced article. Its angle: the decade of progress the world quietly lost in two years, when the pandemic erased gains it had taken since the 2010s to build.
You can read it here and click the 🔍 on any claim to see the data and code underneath it.
So, is it good? You can make your own mind. You get an interactive chart as the opener, where you draw your own guess of how life expectancy evolved before the real line is revealed. Then it digs into the impact of COVID on life expectancy, what changed in the previous decades, and so on. It highlights that 33 years separate Japan (84.46) from Lesotho (51.48) in 2021. A solid analysis, lots of data and visualizations.
But it also reads a lot like AI, with too much information being thrown in the article and some writing tics you can immediately attribute to Claude.
But remember: all I had to do was to copy paste a dataset. That’s it. Not even prompting the model.
This also is an additional demonstration that the artifact is becoming cheap. Building the page - the charts, the layout, the working code - used to be most of the labor. That value is collapsing fast. A polished, publishable, interactive page is now close to free.
How they measured “better”
The researchers recruited 53 readers, blind, and had each one score an AI article and a human one across five dimensions: visual design, narrative and pacing, transparency, claim-data alignment, and insight. The agent came out ahead on all five, and 74% of readers preferred it overall.
And their comparison set is big guns: The Economist and The Pudding.
Does this mean AI has reached a level of editorial sophistication comparable to what’s basically la crème de la crème of journalism?
Before you throw your keyboard to the trash or doubt about your life choices yet, hold on. The reason is about a specific design in the experimentation.
The agent’s single biggest win was auditability. As each claim gets written, the Inspector agent ties it back to the exact artifact that produced it: a line of analysis code, a row in the dataset, a citation. That mapping doesn’t sit in a methods appendix nobody reads; it’s wired into the page through a reviewer tool. Every claim carries a 🔍 you can click to pull up the data and code underneath it, right there in the story. So auditability stops being a promise in the methodology and becomes a button the reader can press.
93% of the AI’s claims linked back to the specific line of code or citation that produced them. For the human-written articles, that number was 25%. It’s not because journalists are sloppy, but because they almost never expose the wiring. The number in a chart is right; you just can’t click it and watch it resolve back to the dataset.
In my life-expectancy story, every figure is clickable back to its origin. No “trust me.” Show me.
That’s almost a side insight of this study, but something we should think about when we design stories.
Also, two caveats before anyone over-reads it: the sample is small (53 readers), and this is a preprint. And the paper is scrupulous about what that means: it measures whether a claim can be traced to its source, not whether the claim is true.
The half it can’t reach
The paper is honest about what the 7 agents team is capable and not capable of. The agents are excellent at exploring what’s in a dataset, and much weaker at finding the story - the angle that comes from outside it.
The researchers measured this too. The agent recovers only about half of a human journalist’s editorial angle. The other half lives in reporting: talking to sources, calling experts, knowing what the numbers don’t say.
Their cleanest example: given a dataset of repair-café records, the AI can rank what breaks most often. It cannot tell you the reason is that manufacturers designed those products to be unfixable. That claim isn’t in the data. It comes from a person who went and asked.
Same with my own test about life expectancy. The agents don’t know what they don’t know. Which raises a question we don’t ask enough: how good are we at giving agents the context they’d need to really collaborate with us?
This also highlights a fundamental nuance.
Visualizing data and doing journalism are not the same job, and this is exactly where they separate.
You can analyze a dataset, pull out a few visualizations and stitch them together. That doesn’t make it data journalism. And I’ll circle back to this: this is where we jump in with curiosity and editorial judgment, so we can co-create with the machines.
Collaborator, not replacement
On the other hand, if you’re thinking “phew, I’m safe, I’m the only one who can talk to a source,” I don’t think that’s a very defensible position either. We can expect AI systems to reach out to sources and ask thoughtful questions in a near future.
The recent Baguette Index is a good example: a team built an agent that called bakeries to compare prices. And with autonomous agents, you give them an inbox, you collaborate with them on what to ask, and we’ll probably see more and more of these gathering capabilities. So again, it’s really a question of how we build collaboration, so we can expand our journalism rather than replace people.
Interestingly, the researchers don’t pitch this as a replacement for journalists:
“We position Data Journalist Agent as a collaborator for human journalists:
agent-generated articles can augment the newsroom workflow by contributing creative multimodal assets and an auditability dimension that is rarely formalised.
Beyond augmenting existing coverage, Data2Story opens a complementary path: surfacing specialised or niche datasets that human journalists rarely have the bandwidth to investigate in depth, turning overlooked data into accessible, verifiable stories.”
So, who’s winning (the World Cup)?
That first aspect is exactly what we’re seeing in our World Cup experiment. What we’re trying to do is surface how AI can connect people with the human creators and voices that actually power the answers. You can read two previews here:
Look at where the interesting angle in each one actually comes from. The France piece isn’t “the favorites will win” (a number could tell you that). It’s the read that a routine three-player rotation is really a referendum on France’s squad depth and a quiet phase-out of a Ballon d’Or winner.
That angle doesn’t live in a stats feed but in French sports radio and the punditry around it. The agents got there by surfacing and weaving together those human voices (L’Équipe, RMC’s Rothen s’enflamme and After Foot, Eurosport).
The Portugal piece is even cleaner. A data point anchors it: 75% possession but a lower xG than the opponent. The line that reframes the whole debate didn’t come from the spreadsheet, but from a DR Congo midfielder saying Ronaldo “no longer demands the same effort from opponents.” That’s the half the agents can’t reach on their own: a source, a quote, a perspective from inside the game.
And the value of the system is that it found that voice and put it at the center of the story.
Which is the whole thesis in miniature: the agents are strong at what’s in the data, and most useful when they’re surfacing and connecting the human reporting that lives outside it.
The more we work with these systems, the more we see how important expertise is. And perhaps more fundamentally, how important it is to design good collaboration mechanisms and interfaces with agents.








Nobody prefers the interview that didn't happen. That is why the 3-to-1 result worries me more than it reassures. Blind readers can only rate the half that made it onto the page, and markets follow what gets measured. The missing half doesn't just need doing, it needs to become visible enough that someone keeps paying for it.
AI is making the artifact cheaper, but not necessarily the insight and judgment behind it.
Charts, layouts, summaries, claim tracing, and even polished interactive pages are becoming easier to produce. But knowing what matters, what is missing, who to ask, and what the data cannot explain still requires human perception and judgment.
The real opportunity may not be replacing the journalist, analyst, or expert. It may be in designing better systems where AI handles the traceable work and human beings bring in the context, curiosity, and editorial judgment.