Predicting human behaviour
Why it is hard, and why LLMs may not help
I listened to a great talk by Dokyun Lee (‘Scylla Ex Machina: The Illusion of LLMs as Human Surrogates’). The ‘human surrogates’ are models which have been suggested by both researchers and businesses as a possible shortcut to making predictions about how humans will react, with obvious promise for scientific understanding, market research and political polling insight. Here’s simile.ai:
“it is possible to simulate real people with high accuracy. We are now developing a foundation model that predicts human behavior in any situation, at any scale.”
Dr Lee gives an overview of the current state of the art on LLMs and what they can (and can’t do). He also rips into the claims that language models are a way of avoiding the effortful gathering of data. His critique is worth listening to in full, but he makes some comments about the fundamental challenges of predicting human behaviour which apply to all of behavioural science, not just the part that might be seeking to use new technology. Because the same limitations hold however you try and predict human behaviour, I thought it was worth expanding on them here.
It is instructive to reduce prediction to a simple form. If the situation is X, what is the predicted outcome, Y? Even this simplest conception gives us a space of possible situations (inputs) and outcomes (outputs).
Science starts by collecting points which have already been observed (evidence) and using them to infer points which haven’t yet been observed (predictions). Both can be expressed as links between situations and outcomes (X-Y pairs).
Then, using the gathered evidence, models are built which capture the available evidence. Good models also capture future evidence (aka predictions). It is the factors which stop models being good predictors which are of interest to Lee (and us).
In the pure world of mathematical functions if you get the right model for the data you know everything about the data - all current observable points as well as all possible future points, out to infinity.
This is the principle that physics has applied so successfully. The same mechanical foundation can predict the trajectory of a ball kicked in New York and that of a rocket fired from Malindi, Kenya. These foundational models, says Lee, are fundamentally different from the synthetic humans created by companies which purport to provide digital twins for individuals or populations using LLMs. Foundational models translate across contexts because they are based on uncovering the generating principles which create the observed data (the “causal model”). LLMs are models which capture the observed data, but don’t encapsulate the underlying principles, so they have no foundation to translate to novel situations.
Why ever might people think these data-models are good at prediction?
Lee observes that LLMs are general purpose ‘function approximators’. They find patterns in data. They can fit the systematic patterns in human text, but they could just as easily find patterns in random noise. When they get trained with human text they compress those data and can so make plausible outputs (recovering patterns, from the simple, like “Happy” is likely to be followed by “Birthday”, to the complex, like the range of responses you get to a prompt when you ask the LLM some complex task).
For powerful function approximators, ‘predicting’ anything which is in the sample is “trivial”. If you have already seen that X1 leads to Y1 in the training data, then predicting X1 leads to Y1 just follows from what you have trained the model to do.
That some of the abilities of LLMs result from this fact is obscured by the fact that LLMs hoover up a vast, incomprehensible, amount of data. This is impressive, but the ability to reflect and re-represent the training data shouldn’t give us confidence that the model can do more than that.
In our schematic, this is as if the model sees points (X1,Y1), (X2, Y2), and (X3,Y3). Questions about X1, X2 or X3 will get perfect “predictions”. Lee is also unimpressed by interpolation within the range of the observed data, something like asking about an input halfway between X1 and X2. Within this scheme, the ability to do both recall and interpolation tells us nothing about the ability to extrapolate beyond the data, to predict, given X4, the output Y4.
Extrapolation is precisely what you want a good model to do. One way to get this is a foundational set of principles which translate across contexts (like Newtonian mechanism does for local space and time). Another is to gather a mass of training data and hope the secret sauce of deep learning will uncover patterns in the data which support extrapolation beyond the data. Lee is sceptical this can be the case for human behaviour.
He is sceptical of the use of LLM models for non-observed data for a couple of reasons.
One reason is the fundamental difficulty of validating a model for cases which you want to make predictions for. By definition the cases which you want to make predictions for are more likely to be novel, and to the extent they are novel there isn’t training data, and if there isn’t data how can you validate? To hear Lee tell it, the whole field of digital twins and synthetic polling using LLMs for human behaviour is snake oil based on assuming the models will produce good outputs when this can’t be known by definition. The outputs are plausible for what is already known, but this strength (capturing observed data) is precisely also their limitation.
The ability to discern this is confused by ‘data contamination’ in the benchmarks used to test models. Because LLMs have hoovered up such a vast amount of training data, many standard tests (and their answers) are included in the training data. When the models look like they are reasoning (extrapolating) they may just be producing the outputs they’ve already seen in their training data (recall).
A second reason that the validation of models of human behaviour is hard, and this is particularly true for LLMs, is that the models are brittle. This means they produce wide variability in their outputs in response to small changes in input. So, for example, small changes in prompt wording —rephrasing that doesn’t change the question — can produce completely opposite answers (indeed, there is a whole area of ‘prompt injection’ dedicated to engineering peculiar ways of phrasing requests that get around the normal response of the language model — usually with the aim of subverting a refusal).
These limitations might be of less concern if humans were Newtonian objects: solid, immutable and affected only by a limited set of physical forces. Instead, though, we are both adaptive (changing based on our experience and the momentary environment) as well as heterogeneous (varying between people, as well as across time). These factors mean that the observable data has limited value in predicting behaviour in new circumstances.
Here Lee’s critique becomes, in my opinion, a general one for psychological science. Modelling observed data will be fundamentally limited without building a causal account, an account of “why” people do what they do. LLMs are incredible new technology, but that doesn’t mean they can escape this fundamental limitation. If they are, at heart, a data-modelling exercise they capture what already is, but not beyond. They demonstrate just how much of the human world has been compressed into language, but they will be limited precisely to what has been experienced and written about so far, and will have limited ability to extrapolate to the truly novel.
If a market researcher wants to know how people will respond to a new product, or a political pollster wants to know how people will vote in an election, the models will give plausible answers. That’s what they do, but the answers will only be plausible, no more. To the extent that the situation were are asking about is novel, exactly when we might gain from accurate predictions, they will be useless.
There’s more in the talk I haven’t covered, it really is an excellent introduction to both the limitations of LLMs and some deeper issues about why psychological science is so hard. Recommended.
This newsletter is free for everyone to read and always will be. If you can afford it, feel free to chip in to help me keep writing for everyone: upgrade to a paid subscription (more on why here).
See below for references, and other things I’ve noticed.
References
Dokyun Lee. Scylla Ex Machina: The Illusion of LLMs as Human Surrogates. 05-07-2026, hosted by the Digital Economy Network https://www.digitalecon.org/seminar
Website of Dokyun Lee: https://www.leedokyun.com/
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337-351.
Other things…
Small announcement
I’m going to take August off, so next week’s newsletter will be the last until September. Happy holidays!
Podcast: The Great Political Fictions: The Dispossessed
From the History of Ideas Podcast with David Runciman, a review of Ursula Le Guin’s novel ‘The Dispossessed’. Runciman identifies Le Guin’s great achievement of portraying an anarchist utopia that seems plausible, both psychologically and politically. He is also draws out that the difference between the anarchist moon colony (the utopia) and the degenerate statists of the origin world is not exactly in freedom from coercion. Freedom exists on both, either by material abundance (origin world), of by the lack of the structured of violent control (anarchist moon), as does coercion, by force (origin world) or convention, upbringing and social pressure (anarchist moon). The real different is the honesty with which the anarchists recognise the compromises they have made to make society work. The corrupt elites of the origin world boast that they are only recognising the real politic of force, but at the same time they give lie to this because they cannot allow the truth about the organisation of their society to be told, it has to be dressed up in sentimental pageantry and ideals which only ring louder and more hollow as they are violated in the streets.
David Runcimen: The Great Political Fictions: The Dispossessed
Catch-up
Social thinking. How other people, even if just imagined, can help puncture the illusion of understanding
Striking evidence for forced experimentation. A natural experiment shows how we always have something to learn, even about the things we are most familiar with.
Another algorithm is possible. How social media could easily be changed to stop us hating each other so much
AI systems out-persuade expert humans. Panic? Quick review of a compelling new report
Why we like to believe other people are stupid. Reason sells, but who’s buying?
… And finally
Funny because true?
END
Comments? Feedback? Synthetic predictions? I am tom@idiolect.org.uk and on Mastodon at @tomstafford@mastodon.online
AI declaration: I write all the words and think all the thoughts myself. I asked Gemini to check for spelling and grammar. For this post I also asked Claude to make the plots (link).








Most of the synthetic personas seem to be used to recreate surveys and focus groups, which are already flawed as prediction devices. The 'say vs. do' gap exists for a lot of conventional market research; recreating it via LLMs is arguably a step in the wrong direction, as you point out.
This is interesting although I'm struggling to take the objections seriously. I believe we already have pretty good predictive models of behavior when it comes to crowd control for example or even more complex things like change management or crime prediction. I fail to see what the "je ne sais quoi" of human behavior might be which would make it impossible to model effectively.
We are a complex system of course, and the maximalist claim is probably false (at least will be for a while) but we already have behavioral models that make reasonably accurate predictions and these objections can't really explain why wouldn't LLMs be able to refine them.