Two AI scribes. One saved time. Clinicians could not tell them apart.

A randomised trial put two ambient AI scribes head to head. One cut documentation time, one cut none, and both made the work feel better. Relief and throughput are not the same purchase.

Share
A clinician in a white coat typing at a desk with a stethoscope beside the keyboard
Photo by Vitaly Gariev on Unsplash

UCLA ran the trial the ambient scribe market had been missing. Three arms, 238 outpatient physicians, 14 specialties, randomised between Microsoft DAX Copilot, Nabla, and carrying on as normal, from 4 November 2024 to 3 January 2025. It came out in NEJM AI last November.

Nabla cut documentation time by 9.5%. DAX cut it by 1.7%, which is another way of saying it didn't cut it. Both groups reported feeling meaningfully better, and DAX produced the larger improvement on exhaustion and task load. Asked directly, the clinicians said the two platforms performed about the same.

So one product moved the clock and the other didn't, both products moved the mood, and the people using them couldn't tell which was which. I've spent a few days working out what a hospital is supposed to do with that, and the answer turns out to be a question I've been getting wrong in my own rollouts for years.

What the ambient AI scribe trial actually found

The numbers, from the published trial, with the open preprint on medRxiv for anyone without access.

The primary outcome was change in log writing time-in-note. Nabla came in at minus 9.5%, confidence interval minus 17.2 to minus 1.8, p equals 0.02. DAX came in at minus 1.7%, interval minus 9.4 to plus 5.9, p equals 0.66. That second interval crosses zero comfortably. DAX did not beat doing nothing.

The secondary outcomes went the other way. On the Mini-Z burnout scale both improved, DAX by 2.83 points and Nabla by 2.69. On physician task load DAX fell 39.9 points with the interval clear of zero, Nabla 31.7 with the interval just touching it. On work exhaustion DAX fell 0.32 with the interval clear, Nabla 0.23 with the interval just touching. On every subjective measure the product that saved no time did at least as well as the one that did.

Two details matter more than they look. Clinically significant inaccuracies were reported as happening occasionally for both, 2.7 and 2.8 on a five point scale, so this is not a story about one product being better made. And adoption was low. DAX was used in 33.5% of 24,696 visits, Nabla in 29.5% of 23,653. Roughly two thirds of the time, a physician who had been given a scribe didn't use it.

Then the same thing happened at 2.33 million encounters

Close up of a mechanical stopwatch on a black background
The clock is the easy thing to measure, which is most of why it gets measured. Photo by William Warby on Unsplash.

One trial with a split result is a curiosity. I went looking for whether it replicated, and it does, at a scale that makes the UCLA trial look like a pilot.

A Spanish health group published a retrospective multicentre analysis of 16 months of outpatient deployment, September 2024 to December 2025, covering more than 2.33 million assisted encounters. Different country, different product, different design, no randomisation.

Adoption climbed from 2.7% of outpatient visits to about 31%. Hold that against UCLA's 33.5% and 29.5%. Three deployments, two continents, all three settling within a few points of one visit in three.

Consultation duration with the scribe was a weighted mean of 15.01 minutes, against 14.65 without it. That is 0.37 minutes longer, not shorter, converging toward parity as the rollout matured. Semantic agreement between what was said and what was written held between 87.4% and 89.2%. Clinician experience improved on five of seven domains, effect sizes between 0.25 and 0.41, consistently across light and heavy users.

Two and a third million encounters. No time saved, arguably a little time lost, and the people doing the work felt better about it anyway.

Three deployments, three answers on the clock. One answer on how it felt.

Change in documentation or consultation time against no scribe. Bars growing right saved time, the bar growing left added it.

Nabla, UCLA randomised trial

9.5% less time

DAX Copilot, UCLA randomised trial, not significant

1.7% less time

Spanish deployment, 2.33 million encounters

2.5% more time

All three reported that clinicians felt better. The Spanish figure is 15.01 minutes per consultation against 14.65 without the scribe. Sources: NEJM AI, November 2025, and the Spanish multicentre analysis of September 2024 to December 2025.

Then a third time, with a cleaner design. A randomised crossover trial published in JAMIA on 1 May 2026 ran two ambient scribes past the same physicians. One product beat the other on documentation time by a clear margin, 3.19 minutes less per day, interval running 4.87 to 1.50 and nowhere near zero. On burnout the authors report that "[b]oth tools reduced personal and work burnout scores, but differences between tools were not meaningful." Same shape a third time. The time number separates the products and the relief number does not.

The reinvestment story, and why I am not ready to believe it

The Spanish authors offer an explanation for their own null result, and they're straightforward about offering it. Their view is that technology should return time to high value clinical interaction and patient relevant outcomes rather than uniformly shortening consultations. The saved effort, on this reading, went into the conversation with the patient, not into the clock.

That might well be true. It's also untested by the data they present, and I want to be careful here, because the same sentence gets said in my industry constantly and it's almost never measured.

When a tool fails to move delivery throughput, the explanation offered is that the time went into quality, or review, or design, or mentoring. Sometimes it did. I've said it myself and believed it. But the sentence has a property that should make anyone uneasy: it arrives after the disappointing number, it converts a null result into a success without any new evidence, and in the form it's usually stated there's nothing that could contradict it. A claim that can absorb any outcome isn't doing any work.

The honest position on the Spanish data is that consultations didn't get shorter, clinicians felt better, and where the difference went is unknown. That's a respectable finding. It just isn't the same as knowing the time was well spent.

What this looks like from inside an engineering team

An empty hospital corridor with benches along the wall and closed doors
Whatever the tool gave back, nothing in the system records where it went. Photo by Tasha Kostyuk on Unsplash.

Here's the version I've lived, and the reason this trial stopped me.

You roll out a tool. Six weeks later the delivery metrics look unchanged. Cycle time flat, throughput flat, the burndown identical to every other quarter. And the team tells you, without being asked, that they would riot if you took it away.

There are two standard readings of that room and both are wrong. The sceptic says the enthusiasm is novelty, the metrics are the truth, cancel the licence. The advocate says the metrics lag, the benefit is real, give it another quarter. Same data, two reasonable people, no resolution. That combination usually means the disagreement isn't about the data.

What the scribe trial makes hard to avoid is that relief and throughput are separate outcomes with separate causes, and a tool can deliver one without the other. Not as a measurement artefact. As a real result. DAX genuinely saved no time and genuinely reduced exhaustion, and both of those are facts about the same eight weeks.

Once you accept that, the rollout question changes shape. It stops being whether it worked and becomes which of the two you bought, and which one you needed. Those have different answers and most organisations have never asked the second out loud.

A team drowning in after-hours admin and a team missing its delivery dates need different things. A tool that fixes exhaustion is a legitimate purchase for the first team even if the clock never moves. What it isn't is a throughput investment, and the trouble starts when it got sold as one internally, because then the null result arrives and somebody has to either kill something that's working or invent the reinvestment story to save it.

The adoption number nobody wants to talk about

The other thing all three deployments share is the ceiling. About one visit in three.

This matters for a reason that's pure measurement. The UCLA trial reported across all visits, which is the right way to run a pragmatic trial, because it tells a health system what it actually gets when it buys the thing. But it means the per use effect is diluted by roughly two thirds. Nabla's 9.5% across all visits is a considerably larger effect inside the visits where it was switched on.

Anyone measuring an internal AI rollout by department wide averages is doing the same arithmetic without noticing. If a third of your engineers use the tool and you measure everyone, you've built a denominator that guarantees a disappointing number, and then you make a licensing decision on it.

Both readings are available, and which one you take is a management decision, not a data one. Pretending otherwise is how these arguments run for two years without moving.

The claim I will defend

Relief and throughput are separate outcomes, a tool can deliver either one alone, and almost every rollout measures one while arguing about the other.

The corollary is the useful part. Before buying, write down which of the two you're buying and how you'll know. If the answer is throughput, the survey doesn't count as evidence however enthusiastic. If the answer is relief, the delivery metric doesn't count as refutation however flat. Most rollouts do the opposite of both, and then the argument runs on vibes for two quarters.

This is a different failure from the one I wrote about last week. That was a cost that moved somewhere the system couldn't see. This is two things the system sees perfectly well, pointing in opposite directions, with no rule for what to do about it.

Prior art, and what is actually new here

The trial is reported accurately elsewhere. Trade coverage carries both figures, DAX at minus 1.7% and not significant, Nabla at minus 9.5%. I checked, because my first instinct was that the split had been buried, and it hadn't been. Nobody hid anything.

What's missing is the reckoning. The headline across the coverage is that scribes may reduce documentation time and burnout, which reads two outcomes as one finding at the exact moment the trial shows them separating.

One roundup does put the two studies side by side, so I am not claiming the juxtaposition. SOAP Note AI's research summary, updated July 2026, carries the UCLA figures and the Spanish ones in the same document and lands on the conclusion that the time saved on typing "gets spent talking with patients instead". Two things are still open after it. That summary logs DAX as a "[s]imilar directional benefit", which turns a result whose interval crossed zero into a smaller version of Nabla's, and the gap between those two readings is the entire argument here. And it treats the reinvestment explanation as the finding instead of a hypothesis, which is what I want to look at next.

The trial authors are careful and say the secondary findings need confirmation in larger multicentre trials. The Spanish paper isn't that confirmation, being observational and not designed to test this. But it is 2.33 million encounters pointing the same way, which is worth more than nothing.

I've made a neighbouring argument before, about what observability actually tells you. The difference is that this one isn't about a metric being wrong. Both metrics here are right.

How this got here, in order

September 2024. The Spanish outpatient deployment begins, scribe adoption at 2.7% of visits.

4 November 2024 to 3 January 2025. The UCLA trial runs, randomising 238 physicians across DAX, Nabla and usual care.

11 July 2025. The UCLA results post as a preprint on medRxiv.

26 November 2025. The trial publishes in NEJM AI, the first randomised evidence on ambient scribes.

December 2025. The Spanish deployment closes its window at more than 2.33 million assisted encounters and roughly 31% adoption.

Through 2026. The trial picks up citations across emergency, inpatient, dental and mental health settings, and ambient documentation starts being described as standard instead of experimental.

Where this stands this month

Ambient documentation is moving out of the consultation and into the rest of the record, which changes the stakes. A note produced during a conversation is one thing. A system generating parts of the record continuously is another, and the accuracy figures in these studies were measured on the easy case: 87% to 89% semantic agreement on a transcript that a clinician reads and signs shortly afterwards. Nobody has published the equivalent number for documentation generated where no one is waiting to check it.

And here's the point against my own argument, which is a decent one. If exhaustion genuinely falls and stays fallen, that's a durable workforce benefit in a system losing people to burnout, and arguing about the clock is arguing about the wrong outcome. Someone could fairly say I'm doing what I accused the advocates of doing, in reverse: demanding a throughput number from a tool that was never a throughput tool. The difference, I think, is that I want it settled before the result arrives rather than after. That's a preference about process, not a proof.

Marking my confidence

Established. Everything from the two papers. The UCLA design, 238 physicians, 14 specialties, the dates, Nabla at minus 9.5% with p equals 0.02, DAX at minus 1.7% with p equals 0.66, the Mini-Z, task load and exhaustion figures, the inaccuracy ratings of 2.7 and 2.8, and adoption of 33.5% and 29.5%. From the Spanish analysis: 16 months, more than 2.33 million encounters, adoption from 2.7% to about 31%, 15.01 minutes against 14.65, semantic agreement of 87.4% to 89.2%, and improvement on five of seven experience domains at effect sizes of 0.25 to 0.41.

Not original. The finding is the trial's. The reinvestment reading of the Spanish null is the Spanish authors' own, stated plainly in their paper. The caution about secondary endpoints is the UCLA authors' caution, not mine.

Inferred. That the same split shows up in engineering tool rollouts, and that the roughly one in three adoption ceiling is a real property and not three coincidences. Three data points is not a pattern I'd bet a budget on.

Guess. That within two years a large health system publicly renews an ambient scribe contract on wellbeing grounds while conceding no measurable time saving, and that it gets reported as a scandal by people who were told it was an efficiency product. Held loosely.

What would change my mind: a well powered trial showing the subjective improvement decays to nothing after a year, which would make it novelty after all, or one showing that among clinicians who actually use these tools consistently the time saving is large and the diluted averages were hiding it.

What I would actually check

Before the pilot starts, write one sentence saying which outcome you're buying, relief or throughput, and what number you'll look at. Then put it somewhere you can't quietly edit later.

Not for rigour's own sake. Because after the result lands, whichever way it lands, there will be a persuasive story available for calling it a success, and by then you won't be able to tell whether you believe the story or just prefer it.