Gerhard Richter, Spiegel, Grau (Mirror Painting, Grey), 1991: a tall grey pane of glass in a dark frame.
Gerhard Richter, Spiegel, Grau, 1991. © Gerhard Richter. National Galleries of Scotland. Photo: Antonia Reeve.

Research in Conversation · MLL Colloquium

What Is Left to Write?

Richter’s Grey, the Machine’s Em Dash — and What Makes Us Us

Gerhard Richter

Anyone can paint a canvas grey.

When Richter painted his canvases grey, it was out of a crisis: he did not know what was left to paint.

The pictures, he said, began to teach him. paraphrase: check in Text (2009)

This one is a mirror painting: Spiegel, Grau, 1991, grey pigment on glass.

Gerhard Richter, Spiegel, Grau (Mirror Painting, Grey), 1991.

Many in the academy know the feeling now

Is the machine good? For me, that is settled.

A proof

“The argument in this paper was entirely found by ChatGPT Pro 5.5.”

A question open since 1996 (Benjamini and Schramm); four pages, May 2026; unrefereed pending

“short and particularly elegant”

Hugo Duminil-Copin, 30 August 2026

A map

A problem that had thwarted her group for years; ChatGPT resolved it in a few minutes.

Alyssa Goodman’s spiral-arm map, as reported in Science, 4 June 2026 pending

A playlist

People I know listen mostly to AI-generated music.

97% could not pick out the AI track in a blind test; 80% want it labelled.

Deezer and Ipsos survey, 9,000 adults, October 2025 (vendor figures) pending

Placeholder · the cost

More than we can check

Understanding: Terence Tao, Scientific American, 10 September 2026: “We are very, very close to a scenario in which a major result gets proved, and no human can understand and explain it”; answers and insight have become “negatively correlated” (§12.10). The flood: Bloom, Science News, 8 June 2026: “hundreds of pages of math that they can’t understand or even read … Who’s going to be able to check this?” (§12.10); in code, “AI slop is cheap to generate and expensive to review” (Baltes, Cheong and Treude, preprint). Spoken only, nobody named: the promotional email. For the panel, not a claim: would we still recommend studying mathematics, computer science, translation? (Overheard answer for mathematics: as much as before.)

Two halves of one college

The sciences

Does the work check? A proof checks, or it does not.

The arts

Creation, mediation, reception: a novel is written, carried, and read.


“Anyone working in astrophysics is someone who wants to do astrophysics, not someone who wants to learn the answers.”

David Hogg, quoted in Science, 4 June 2026 pending

What is left to write?

A question for you, and for the conversation.

—

What I can do is measure one mark.

One mark, all the way round

Placeholder · human convention

The mark was always ours

Dickinson in one line (her editorial history to verify first); human use varies about 45 times across registers (§2.5); newspapers have house styles too (work in progress).

Model behaviour: the rise and the return

The flagship came back to the human rate, indistinguishable from it; the small tier went the other way.

In science, where human writers barely use it

GPT-5.4: 2 and 1 dashes in about 23,600 tokens, inside the grey. Consistent with a deliberate correction; intent is unprovable. counts pending

Placeholder · the tell moves

Delve goes out, the dash comes in

Tell-migration, correlational, said out loud (§1.8, deep-N tier: never on a slide with uniform-tier numbers). The Economist, 30 July 2026: “Today only Claude uses more em-dashes than human writers” (p. 4; §8.2 to be made verbatim).

The reference problem

More than which humans?

So “em dash, therefore AI” has no denominator. It is not a weak test; it is not a test.

Why? Mechanisms

Per million tokens · science continuations · 42,000 items per model

Tokenizer? Ruled out: GPT-4o and GPT-4.1 share one and differ about 30 times.

Preference training? The recipe decides: four families up, two down.

Rewarded format. Preference models favour lists, links, bold text and emojis (Zhang et al., ACL 2025). The em dash may be one more.

Markdown? A preprint’s idea: testable, and we are building the test. work in progress

Where does it come from? One open model, stage by stage

In the post-training data, the dashes come with ChatGPT’s answers, and preferred answers carry fewer than rejected ones. work in progress

OLMo-2 (Ai2) publishes every stage’s training data; counted by us, 29 September 2026: pretraining from 93 sampled shards (a whole-corpus count agrees within 1%), post-training in full. About a tenth of the chat data’s dashes are the Chinese double dash.

The same stages, in one model’s own writing: OLMo-2 1B

Writing freely, it starts below people; the preference step takes it past them. work in progress

OLMo-2 1B (Ai2), four released checkpoints, each continuing the same 10,000 news articles (first halves), sampling from its own probabilities, first 40 words; whiskers are 95% intervals. Always taking its likeliest next word instead, it writes a tenth as many. Counted by us, 29 and 30 September 2026.

The same writing, closer up: the dash’s form

People writing news space the dash; after training, the model mostly closes it up. work in progress

Share of em dashes closed up rather than spaced, first 40 words: the same 10,000 news articles, and OLMo-2 1B’s continuations of them, sampled from its own probabilities. Counted by us, 30 September 2026.

Placeholder · the trace

House styles, not a language

Within one 2026 cohort, contractions range from about 1,200 to over 30,000 per million (§4.1): Rudnicka’s idiolects. “AI language” is the wrong noun.

From heuristic to accusation

7,806 entries to one prize

At a 0.5% false-positive rate, about 39 innocent stories are flagged.

With 2% AI entries and 95% caught, one flag in five is wrong.

At 1% and 1%: a coin flip.

The flag is not the verdict

Two prizes, two flags

Commonwealth Short Story Prize, 2026
Goncourt longlist, September 2026 disputed
The flag
A detector: 100% AI.
One detector: over 95% AI; others: likely human. The author denies it. pending
What followed
Drafts, time-stamped documents and notes examined.
Plagiarism turns up in his journalism of the 2010s; the novel is dropped from the list.
Where it stands
Cleared on 22 June; the overall prize on 30 June.
FAZ: „Seitdem steht der Autor mit dem Rücken zur Wand.“

„… wer das Buch liest, kann sich nur wundern über dessen Fürsprecher.“

Niklas Bender, Frankfurter Allgemeine Zeitung, 28 September 2026. On the prose, AI or not: „komplett missraten“.

The loop closes

Writers strip the mark

“When I write ‘delve,’ I change it. It has this stigma now.”

Zina Ward, co-author of the delve paper, in Laura Yuen, Minnesota Star Tribune, 4 December 2025

My own grant documents follow a zero-em-dash convention. On purpose.

Do we pick up the machine’s words? In unscripted speech, they rise; the cause is open (Anderson, Galpin and Juzek, AIES 2025). Association, not causation.

Beyond the mark: meaning

“Does money lead to happiness?”

About 100 US adults, one essay each; GPT-4o judged each essay’s stance.

Without AI, two in five essays took no side; when the AI wrote much of it, two in three.

In a second test, asked only to fix the grammar, the models still moved the meaning.

And what the machines write, we all read.

Abdulhai, White, Wan, Qureshi, Leibo, Kleiman-Weiner and Jaques, “How LLMs Distort Our Written Language”, arXiv 2603.18161, v2 August 2026 (preprint): Figure 6, redrawn. The AI assistant was assigned at random; how much to use it, people chose. The grammar test: 86 essays written in 2021.

Circling back

Can alignment be achieved, even in literature?

Literary quality is measurable, and crowd and expert judgements come apart (Bizzoni et al. 2023; Feldkamp et al. 2024).

What is left to write?

Over to the conversation.

With Joachim Adams and Susan Bryson

Gerhard Richter, Spiegel, Grau (Mirror Painting, Grey), 1991.

How these slides were made

Tommie Juzek set the argument, the running order and the substance; the numbers come from his group’s research data and the cited sources, and any not yet logged in his evidence file are tagged pending. Claude (Anthropic’s Opus) built this page from his notes and direction, drew the figures from those numbers, and drafted the slide text. draft v0.3, before Tommie’s edit


Image: Gerhard Richter, Spiegel, Grau [Mirror Painting (Grey)], 1991, pigment on glass. © Gerhard Richter. ARTIST ROOMS National Galleries of Scotland and Tate. Photo: Antonia Reeve. nationalgalleries.org/art-and-artists/89110

Type: Newsreader and Inter (SIL Open Font License). Slides: reveal.js (MIT). Numbers and sources: the two slides after this one.

Every number on these slides

NumberWhatSource
1,078em dashes per million tokens, human news (CC-News)§2.2
0.0human news (WMT): 0 in 91,317 tokens, below 33 per million§2.1
0 in 3.5 MWMT again at 40 times the depth (work in progress)02 pilot chunk 1
3,407 to 4,161GPT-4.1, news, two runs: 3.2 to 3.9 times human§1.2
924 and 837GPT-5.4, news: 0.8 to 0.9 times human§1.5
4,203 and 4,553GPT-5.4-nano, news: 3.9 to 4.2 times human§1.6
2,077 to 2,288GPT-4.1, science; matched human science 0§1.3
42 to 85GPT-5.4, science: 1 and 2 dashes in about 23,600 tokens§1.4 (counts pending)
4 up, 2 downbase to instruction-tuned, six open families§3.3 to §3.4
491OLMo-2’s pretraining web text (work in progress)§16.2
185; 1,032its instruction data: answers by people; from ChatGPT (work in progress)§16.4
411; 628its preference data: preferred; rejected answers (work in progress)§16.5
679; 513; 1,409; 1,160OLMo-2 1B continuing news, sampled: base, SFT, DPO, final (work in progress)§16.12
966the same 10,000 news articles’ own writers (work in progress)§16.10
13%; 35 to 80%em dashes closed up: those writers; OLMo-2 1B, base to final (work in progress)§16.13
373; 412the 1B’s instruction answers; its preferred answers (work in progress)§16.7
7,806; 39prize entries; innocent flags at 0.5% false positives§5.1; §7
80%; 49%share of flags that are right: 2% prior, 95%, 0.5%; 1% and 1%§7
97%; 80%could not tell the AI track; want AI music labelled (vendor)pending row
32.6, 39.5, 27.9essays for, neutral, against: written without AI (preprint)§13.5
31.0, 44.8, 24.1the same, with light AI use: advice, look-ups (preprint)§13.5
22.2, 66.7, 11.1the same, with the AI writing much of the essay (preprint)§13.5
86essays from 2021 in the grammar-only test (preprint)§13.6

Sources

Em dash measurements: Juzek et al., work in preparation (projects 02 and 07; NSF CAREER proof of concept). Idiolects: Rudnicka and Juzek, arXiv 2608.06589. Uptake in speech: Anderson, Galpin and Juzek, AIES 2025.

Prize: Commonwealth Foundation statements of 19 May, 22 June and 30 June 2026. Base rates: worked arithmetic with illustrative parameters (a 2% or 1% prior, 95% caught, 0.5% or 1% false positives).

Format bias: Zhang, Xiong, Chen, Zhou, Huang and Zhang, “From Lists to Emojis”, ACL 2025, pp. 26940 to 26961. Literary quality: Bizzoni et al., NoDaLiDa 2023; Feldkamp et al., JCLS 2024.

Percolation: “Isoperimetric dimension > 1 implies pc < 1”, 28 May 2026; Duminil-Copin, “Care for a little more AI?”, 30 August 2026. Astronomy: J. Sokol, Science, 4 June 2026. Music: Deezer newsroom, 21 July 2026; Deezer and Ipsos, November 2025.

Stage by stage: Ai2’s published OLMo-2 training data (olmo-mix-1124, dolmino-mix-1124, the Tulu 3 SFT mix for OLMo-2, the OLMo-2 7B preference mix), counted by us on 29 September 2026 (work in progress); the whole-corpus check with infini-gram; the OLMo-2 1B checkpoints (base, SFT, DPO, instruct, April 2025 release) and their own post-training mixes, continuing 10,000 CC-News articles, sampled and greedy (30 September 2026).

Stance: Abdulhai, White, Wan, Qureshi, Leibo, Kleiman-Weiner and Jaques, “How LLMs Distort Our Written Language”, arXiv 2603.18161v2, 26 August 2026 (preprint). Goncourt: Niklas Bender, „Feigheit vor dem Thema?“, Frankfurter Allgemeine Zeitung, 28 September 2026; the detector figures and the denial still to be sourced.