Thesis
We solved hearing the words. Not what they point at.
Nobody designed the keyboard around how people talk. It came from the typewriter, and we adapted. People speak about three times faster than they type, and make about a fifth fewer mistakes doing it.1 Around half of adults cannot write the kind of careful prose a good prompt asks for, though every one of them can say what they want out loud.2 And for a lot of work the keyboard is not slow — it is absent. A surgeon has both hands busy, and so does an engineer inside a machine, a nurse on a ward round, a mechanic, a driver. Others are already talking anyway: a doctor in consultation, a lawyer taking a statement, someone on a sales call. For all of them voice is not competing with a keyboard. It is competing with the work not getting written down at all.
But the models were never trained on anyone talking. Almost everything they learned is written text from the internet — things people sat down and composed — and nobody writes the way they speak. So talking to a model the way you would talk to a colleague makes it worse, not better: across five models and 550,000 scored answers, spoken phrasing cost about ten points of accuracy where typing mistakes cost three.3 The obvious fix is worse still. Tidying the speech into clean written English before the model sees it cost twenty-four points — the most damaging thing in the whole study — and deleting the ums recovers nothing.4 That is the clue. When you speak, the person listening is already in the situation with you, so you point at things instead of describing them — the launch, the 14th, the three teams. Linguists call this exophora: words whose meaning sits outside the sentence, in the room. The sharpest of them are deictic — here, now, this, tomorrow — words that carry no fixed meaning at all, only a position relative to whoever is speaking and when. An agent occupies none of those positions. A model is probabilistic, so it cannot hold that gap open: it settles it. What you were pointing at quietly becomes whatever that phrase most often means. Writing spells things out instead: the meaning sits inside the text, not in the room. That is endophora, and writing has no choice about it — the reader is somewhere else, and an agent is always somewhere else. That is why typing to one works at all: you were already writing for something that could not see what you meant. Speech never had to, so it does not. Speak to an agent and you hand it a sentence built for someone who was in the room. The words arrive. What they were pointing at stays behind, and the model has no habit of asking. That part is fixable. An agent that asks when it is unsure finished 69% of vague requests, against 71% for one handed the full specification.5
The same gap opens a second time in every language that is not English, and in much of the world it opens wider. Arabic, Swiss German and Tamil are diglossic: the written variety and the spoken one differ enough to be learned separately, and it is the written one that is on the internet in quantity. There, “nobody writes the way they speak” stops being a tendency and becomes the arrangement — a model can have read a language thoroughly and still have read almost nothing anybody says out loud. Recognition is not the problem there either — the models read hundreds of them — but the systems underneath are named in English: the files, the addresses, the fields, the statuses. So a request spoken in Hindi or Spanish has to reach the same objects as one typed in English, and the words doing the pointing are exactly the ones a translation step flattens on the way. Spanish su marks neither the owner’s gender nor their number: su calendario is his, hers, theirs or yours, and nothing in the sentence says which one is about to be shared. Hindi उनको is one word for a single person addressed with respect and for several at once, and the verb agrees either way, so nothing says whether one address goes on the message or three. Neither is a translation problem. No amount of language gets you there; only looking does. Typed and exact, spoken in English, spoken in your own language: however the request comes in, it has to end at one outcome or it was never one interface. That is what we are building: a layer that keeps hold of what your words were pointing at, and a harness that carries a spoken request all the way to the action. The keyboard is not going anywhere — it is still the best way to say something carefully, in your own time. Voice becomes the one you reach for first — and agents reach the work that was never going to be typed: the operating room, the ambulance, the rig, the deposition.
Nobody has measured the one combination that matters: natural speech, into a system that holds the context and asks before it acts. Every voice study feeds a system that cannot ask. Every clarification study types. That gap is the experiment we are running.
Sources
- 1
Ruan, Wobbrock, Liou, Ng & Landay — Speech Is 3x Faster than Typing · Stanford / Baidu / UW, 2016
Speech input was 3.0× faster than a smartphone keyboard for English, with 20.4% fewer errors. Measured on people reading text out; the advantage is smaller when you are composing.
- 2
Nielsen — The Articulation Barrier · 2023, updated 2026
On OECD literacy data, about half the adults in rich countries fall below the level a good written prompt assumes. Nielsen's own estimate of who writes well enough for advanced use is 10–20%.
- 3
Hu, Segura, Rostami & Thomason — Should We Type or Talk to LLM Agents? · USC / ISI, August 2026
Five instruction-tuned models, six benchmarks, 550,000 scored generations. Spoken phrasing cost 9.7 points of accuracy; keyboard typos cost 3.0. Casual speech cost 7.4 against 4.1 for prepared speech — the closer to written form, the less it hurts.
- 4
The same study, on cleaning speech up first · USC / ISI, August 2026
Rewriting a spoken request into concise written form cost 24.1 points, the largest effect in the whole suite. Deleting the fillers from a spoken transcript recovered −0.22 points — nothing.
- 5
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents · 2026
An agent that asks when it is unsure resolved 69.40% of underspecified tasks, against 70.80% for an agent handed the fully specified issue. It asked more on hard tasks and held back on easy ones.
Cite this
Rovers. (2026). We solved hearing the words. Not what they point at. Rovers Research. https://rovers.sh/thesis
@misc{rovers2026thesis,
title = {We solved hearing the words. Not what they point at},
author = {{Rovers}},
year = {2026},
url = {https://rovers.sh/thesis}
}