Currently
Thomas Stephan Juzek
Computational linguist · Florida State University
A running note on what I am working on, reading, and excited by right now. I hope to update this page from time to time.
Last updated: August 2026
Right now
Working on
For a recent NYT piece, Vauhini Vara asked me about the what and why of em dash usage in chat models: how much do LLMs use them, and why is it so out of line with human writing? Digging into data from earlier word-choice projects, em dashes turned out to be far messier than word choices. With words, models heavily overuse a fixed set, and this overuse stems to a good degree from post-training. With em dashes the behaviour depends heavily on the model family, and the roles of the pre-trained model and of developer choices are still open questions. What is clear is that ChatGPT's em dash usage has climbed sharply across generations, so I am following up on this.
Recently out
New preprint with Karolina Rudnicka: Beyond "AI Language": The case for the idiolectal nature of LLM output.
Reading
Just out of reviewing for two journals, and about to review for a really nice workshop. Beyond that, trying to keep up with a field that produces a lot of output.
Thinking about
I recently read Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy by Adarsh Kumarappan and Ananya Mujoo: it's impressive technical work, and it contributes to a nuanced (I know, I know) picture of how pretraining, post-training, and context interact in producing sycophancy.