7 Comments
User's avatar
Heather Myers's avatar

Most of the synthetic personas seem to be used to recreate surveys and focus groups, which are already flawed as prediction devices. The 'say vs. do' gap exists for a lot of conventional market research; recreating it via LLMs is arguably a step in the wrong direction, as you point out.

Peter Barrett's avatar

This is interesting although I'm struggling to take the objections seriously. I believe we already have pretty good predictive models of behavior when it comes to crowd control for example or even more complex things like change management or crime prediction. I fail to see what the "je ne sais quoi" of human behavior might be which would make it impossible to model effectively.

We are a complex system of course, and the maximalist claim is probably false (at least will be for a while) but we already have behavioral models that make reasonably accurate predictions and these objections can't really explain why wouldn't LLMs be able to refine them.

Tom Stafford's avatar

Well I think there are a few things here

- I would say that the areas where there are good predictive models are ones where human behaviour is overdetermined (e.g. in crowd situations were tend to act like particles, because how we feel doesn't really bear on our limited range of possible actions). Are there counter examples?

- I agree that there is no "je ne sais quoi" which means humans couldn't be modelled effectively. This should be the whole ambition!

- I also agree that LLMs could be involved in refining existing models, and improving the predictions they provide

- However, Dr Lee's talk was about using base LLMs (with some fine tuning and/or personnas) to predict human opinions and behaviours. In this case, there is a complex, inscrutable, data model and all the issues of a) model instability/bittleness and b) out of sample validation do hold. That, I think, is the core of the argument.

Peter Barrett's avatar

I guess the core of the argument is an empirical question, so I'm looking forward to seeing how brittle these models really are. My understanding is that we are capable at the moment to create highly predictive models of both very micro behaviors (visual attention, moto movement, etc.) and large scale behavior (crowd, population mobility, consumer demand, language production, etc.). With more structured data on individual preferences and behaviors , we should be able to make more accurate predictions about individuals. For example, personality inventories will probably improve with LLMs.

I'm looking forward to how dystopian all of this gets

Tom Stafford's avatar

Our intuitions are different here. I don't think personality inventories have much headroom for improvement. My hunch is that they capture something about the language we use to talk about ourselves (which the models will also be in hock to), but they don't align closely with the stack of causal mechanisms which actually generate behaviour. That's only a hunch though

We probably do agree that there's significant risk of dystopias!

Mike X Cohen, PhD's avatar

Great post as always, Tom. It's very common in psychology studies to average over large samples, and with so much individual variability, the average is not actually representative of any individual. So I think that some of the criticisms your cite could also apply to over-interpreting averages of many humans.

Not that I want to replace humans with LLMs (in research or any other area), but I wonder whether the average of many LLM personas is really worse than (or even different from) the average of many humans? I wouldn't trust an LLM reproduction of *one individual's* reactions, but is the average of 100 LLM personas different from the average of 100 humans?

Tom Stafford's avatar

Absolutely, the criticism applies to all over interpreting all models (and using a population average to predict individual behaviour is a model, albeit a thin one)

Dr Lee also makes the same point you do, that it is easier to capture the population average than individual behaviour. He provides evidence that a) LLM surrogates are often evaluated on producing population averages only, and terrible at capturing either dispersal information (variance) or realistic individual behaviour. I didn't verify the references for this.

My take away would be that we should make a fair comparison of all models of human behaviour, and that should involve a) out of sample validation (ie tests the models haven't seen before) and b) probably some common benchmarks. Dr Lee argues that LLMs are massively overhyped and fail on both of these