Blog10 min
The one from yesterday.
A colleague needs eight words. A prompt box needs fifty. Conversation gets from one to the other in six passes, and the average conversation with a model is two turns long. The cost is not the typing. It is the requests you stop making.
Someone puts their head round the door and says: can you send the one from yesterday to Maya? You know which one. You send it. The whole exchange costs eight words and about four seconds, and neither of you thinks about it again.
Now type that same request into a prompt box. The one from yesterday is not a phrase you can use, because nothing on the other side was there yesterday. So you name the file. You say which version, and which Maya, and what you want done to it. By the time the request is unambiguous it is fifty words long, and you have spent more effort describing the thing than the thing was worth.
That gap is the tax everyone building on top of these systems is paying, and almost nobody is measuring it. It has been measured before, though — carefully, forty years ago, on a task nobody would think to connect to this one.
What that phrase is doing
Halliday and Hasan gave the distinction its name in 1976. Reference that reaches outward, into the situation the speakers are standing in, is exophoric. Reference that reaches back into the text is endophoric. Both are instructions to go and retrieve something from elsewhere, but only one of them can be satisfied by reading.1
The one from yesterday is exophoric. It does not point at any earlier sentence in the conversation, because there was no earlier sentence. It points at a shared afternoon.
Chafe explained why that is cheap. Information the speaker takes to be already active in the addressee’s consciousness gets produced in an attenuated form — lower pitch, weaker stress, pronominalised, shortened. The shortening is not laziness or sloppiness. It is a claim about what is currently in the other person’s head, and in conversation it is overwhelmingly right.2
Which means the eight-word version is not a compressed form of the fifty-word version. It is a different act. Fifty words describe a document. Eight words point at one, and the pointing works only because you are both standing in the same afternoon.
Somebody put a price on it
In 1986 Clark and Wilkes-Gibbs sat eight pairs of students on either side of an opaque screen. Each had twelve cards showing tangram figures — abstract arrangements of geometric shapes, deliberately chosen because they have no name. One partner, the director, had them in a target order. The other had the same twelve in a random one. The director’s job was to talk the matcher into the same arrangement. Then they did it again, and again, six times, with the figures reshuffled between each pass. The transcripts run to 9,792 words across 576 placements.3
Here is one director on one figure, all six times.
Twenty words down to three, and the third word is doing all the work. Across the whole study the pattern holds: directors used an average of 41 words per figure on trial 1 and 8 on trial 6, F(1,35) = 44.31, p < .001. Turns per figure fell from 3.7 to about one, F(1,35) = 79.59, p < .001. The steepest drop was between trial 1 and trial 2, and by trial 6 there was almost nothing left to save.3
Something else changed at the same time. On trial 1 the directors always described the figures, usually indefinitely: a person who’s ice skating. From trial 2 on, 89% of their opening utterances were identifications with definite reference: the ice skater. They had stopped telling the matcher what the figure looked like and started pointing at a thing they both already had a name for.3
The paper’s explanation is the principle it is famous for. In conversation, participants try to minimise their collaborative effort — the work that both of them do from the start of a contribution to its mutual acceptance — rather than each minimising their own. And that produces a trade-off the authors state plainly: the more effort a speaker puts into the initial noun phrase, the less refashioning it is likely to need.34
So the 41 words on trial 1 were never the price of describing a tangram. They were the price of describing one to somebody with whom you had not yet agreed on a way of referring to it. The eight words on trial 6 are what that same figure costs once you have.
What made the eight words possible
Five years later Clark and Brennan set out what a medium has to afford before any of this can happen. Eight constraints: copresence, visibility, audibility, cotemporality, simultaneity, sequentiality, reviewability, revisability. Take one away and the costs of grounding redistribute — they enumerate eleven of those, from formulation and production through to fault and repair — and people change technique accordingly.4
The first column is the one this post is about. Copresence is what licenses the one from yesterday: it is the affordance that makes exophora available at all, because exophoric reference points into a situation and there has to be a situation that both parties are in.
A prompt box has three of the eight. It is reviewable, it is revisable, and turns cannot get out of sequence the way they can in a mail thread — that last one is our call rather than theirs, and it is the only cell in the figure we had to reason about instead of read off. It has no copresence, no visibility, no audibility, no cotemporality, no simultaneity. On their scale it is somewhere between a letter and an email.
And they predicted exactly what you do about that, in 1991, in a paragraph about postal mail: in media that are not cotemporal, repairs made by the other party become very costly, so speakers try hard to avoid relying on anyone else to repair a misunderstanding — it is less costly to revise what you say before sending it. Elsewhere, more bluntly: to avoid paying fault costs, speakers may elect to pay more in formulation costs.4
That is a description of writing a prompt, published thirty-five years before anybody had to write one.
Trial one, forever
Here is what makes the prompt box worse than the letter it resembles. A letter is addressed to a person, and that person accumulates common ground with you across letters. The tangram directors got six passes at the same twelve figures with the same partner, which is what bought them the ice skater.
LMSYS-Chat-1M is a million real conversations, 25 models, 210,479 users, 154 languages. The average conversation is 2.0 turns long. The average prompt is 69.5 tokens by Llama-2’s tokenizer. Chatbot Arena in the same table runs 1.2 turns and 52.3 tokens; Anthropic’s helpfulness set, 2.3 turns and 18.9 tokens.5
Two turns is one request and one reply. Nothing in it refers to the same thing twice, which means the process that turned twenty words into three never starts. There is no trial 2. There is certainly no trial 6.
Those corpora skew towards people evaluating models rather than people getting work done, and we would not lean on the exact figures. But the shape is not subtle and it is the same in all three: whatever the median interaction is, it ends long before convergence would have had a chance to begin. You are permanently on trial 1, and 69.5 tokens is roughly what trial 1 cost the tangram directors — for an abstract shape, to a stranger behind a screen.
Why over-specifying is the correct move
It would be easy to read all this as a criticism of the way people write prompts. It is the opposite. Given the medium, over-specifying is rational, and the numbers say how rational.
Yang and colleagues measured what happens to the requirements you leave out. Models supply an unstated requirement by default 41.1% of the time. Which means nearly six times in ten, something you did not say is not supplied. Worse, the successes are unstable: underspecified prompts are twice as likely to regress when the model or the prompt changes, sometimes with accuracy drops over 20 points. And simply specifying everything does not reliably rescue it either, because instruction-following is finite and requirements conflict.6
Read that as a reference-failure rate and the medium looks worse than a letter, not better. A letter that fails to refer comes back with a puzzled question. A prompt that fails to refer comes back with a confident, complete, well-formatted answer to a slightly different request. Clark and Brennan’s fault costs assume the fault is detectable. Here it frequently is not, and if the request had side effects it is not recoverable either.
So you pay formulation cost to avoid a fault cost you cannot see coming and cannot undo. Everybody does this. Everybody is right to.
The part that costs more than the typing
Here is where we think the trade-off framing stops early. Least collaborative effort is not only a theory about how you say a thing. It is also, unavoidably, a theory about whether.
Every request has a value and a price. When the price is eight words, an enormous number of small requests clear the bar. Is this the latest one. Does Maya still need it. Check the date on that. Those are most of what a colleague actually does for you over a day, and not one of them survives a fifty-word toll.
You notice the tax on the large requests. It is irritating, it is slow, you complain about it. You do not notice it on the small ones, because a request you decided not to make leaves nothing behind. There is no abandoned draft, no half-typed prompt, no error. You simply thought better of it, in about a quarter of a second, and moved on.
Which is the claim we want to put on the record: the cost of reference does not merely make your requests longer. It censors which requests exist. What reaches the model is not a sample of what you wanted done. It is the subset that was worth fifty words.
That is a worse problem than latency, because every measurement built on prompt logs inherits it. Analyse a million conversations and you learn, precisely and at scale, what people were willing to pay fifty words for. You learn nothing at all about what they would have asked at eight. The distribution was filtered before it was recorded, and the filter is invisible in the data because it operated upstream of it.
Subramonyam and colleagues named a neighbouring problem: the gulf of envisioning, where users do not know what the task should be, how to instruct the model, or what to expect back.7 That is a knowledge gap. This is a price gap. They compound, and they are not the same thing — you can know exactly what you want, know exactly how to ask for it, and still not ask.
What voice changes, and what it does not
The obvious move from here is speech, and the obvious argument is production cost: talking is cheaper than typing, and Clark and Brennan say so directly. That is true and it is the least interesting thing about it.
What speech restores from the eight constraints is audibility, and it makes cotemporality possible. What it does not restore, on its own, is copresence. Say the one from yesterday into a microphone and a modern recogniser will transcribe it perfectly — every word, correct, first time. It will still refer to nothing. The transcript is not the problem and never was.
That is the actual engineering problem and it is the one we are working on. Exophora is cheap only when the other side is standing in the same situation, so making it cheap means resolving those spans against the situation — the calendar, the files, the people, the day it was said — rather than asking the speaker to eliminate them first.
And there is a trap on the way there worth naming, because it looks like the solution. The tempting shortcut is to have a model expand the one from yesterday into the fifty-word version and hand that to the agent. That is the same operation we measured costing 24.1 points of downstream accuracy in the last post: the rewrite that reads best is the one that loses the reference.8
What we have not shown
We have not measured the shrink. The claim in figure 4 follows from a forty-year-old cost model and two current corpora, and it is a prediction, not a result. Measuring it means counting requests that were never made, which is exactly what no log contains. We think it is tractable — give matched groups the same work with and without a resolver, and count requests issued rather than requests completed — and we will publish it whichever way it comes out.
The unit comparisons are loose and we have kept them apart in the figures for that reason. A word is not a Llama-2 token. A turn spent on one tangram figure is not a turn spent on a whole conversation. What is comparable across the two is the number of passes over the same referent, which is why that is the only thing figure 3 counts.
Clark and Wilkes-Gibbs is eight pairs on one task, and the tangram figures were chosen precisely because naming them is hard. Most work is not tangrams. The direction of the effect has replicated widely; the size of it, in an office, on a Tuesday, has not been established by us or by anyone we would cite.
The one-sentence version is this. Your colleague is cheap to talk to because they were there yesterday. Everything we are building is an attempt to make that true of software, and the measure of whether it worked is not how well it understands the fifty words. It is whether you ever have to write them.
Sources
- 1
Halliday & Hasan — Cohesion in English — Longman, 1976
Names the distinction the whole post turns on. Reference that reaches out of the text into the situation is exophoric; reference that reaches back into the text is endophoric. Both are instructions to retrieve information from elsewhere, but only one of them can be satisfied by reading.
- 2
Chafe — Givenness, contrastiveness, definiteness, subjects, topics, and point of view — in Li (ed.), Subject and Topic, Academic Press, 1976
Given information is what the speaker assumes is already in the addressee's consciousness, and it is spoken in an attenuated form: lower pitch, weaker stress, pronominalised, shortened. Chafe's later Discourse, Consciousness, and Time (1994) splits the same idea three ways — active, semi-active, inactive.
- 3
Clark & Wilkes-Gibbs — Referring as a collaborative process — Cognition 22(1), 1986, pp. 1-39
Eight pairs of students, twelve tangram figures, six trials, 576 placements, 9,792 transcribed words. Directors used an average of 41 words per figure on trial 1 and 8 on trial 6, F(1,35) = 44.31, p < .001; turns per figure fell from 3.7 to about one, F(1,35) = 79.59, p < .001. The paper proposes the principle of least collaborative effort.
- 4
Clark & Brennan — Grounding in communication — in Resnick, Levine & Teasley (eds.), Perspectives on Socially Shared Cognition, APA, 1991
Eight constraints a medium may impose — copresence, visibility, audibility, cotemporality, simultaneity, sequentiality, reviewability, revisability — and eleven costs that move when one of them is missing. Table 1 scores seven media against the eight.
- 5
Zheng et al. — LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset — ICLR 2024
One million conversations, 25 models, 210,479 users, 154 languages. Table 1: 2.0 turns per sample and 69.5 tokens per prompt, counted with Llama-2's tokenizer. The same table gives Chatbot Arena 1.2 turns and 52.3 tokens, and Anthropic HH 2.3 turns and 18.9 tokens.
- 6
Yang, Shi, Ma, Liu, Kästner & Wu — What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts — ACL Findings 2026
Models infer unspecified requirements by default 41.1% of the time. Underspecified prompts are twice as likely to regress across a model or prompt change, sometimes with accuracy drops exceeding 20%. Specifying everything does not reliably fix it either: instruction-following is limited and requirements conflict.
- 7
Subramonyam, Pea, Pondoc, Agrawala & Seifert — Bridging the Gulf of Envisioning — CHI 2024
Names the neighbouring problem: users do not know what the task should be, how to instruct the model to do it, or what to expect back. That is a knowledge gap rather than a price gap, and the two compound.
- 8
Hu, Segura, Rostami & Thomason — Should We Type or Talk to LLM Agents? — USC / ISI, August 2026
The source behind the 24.1-point figure quoted here. Rewriting a spoken request into clean written English was the most damaging single operation across five models, six benchmarks and 550,000 scored generations.