Part IV
Synthetic audience research: what it can and cannot tell you
I like synthetic research. I also think we are going to hear quite a lot of nonsense about what it supposedly proves. Those positions are perfectly compatible.
If I have a hypothesis, idea, message or question that I might take into a focus group, I am happy exploring it with a synthetic audience first. I can ask what they think, ask why, follow an unexpected answer, change the stimulus and continue for as long as the conversation remains useful. That is an extraordinary capability. It is not the same as having spoken to real people.
Start with what it is
A synthetic audience is an attempt by a model to respond from the perspective of a defined group. The quality of the simulation depends on how well the audience is defined, what the model knows about the category, which evidence has been provided and which assumptions have been made.
You are not interviewing thousands of consumers hidden inside the model. You are asking a model to simulate how those people might respond. The distinction tells you what kind of evidence you have.
Treat it as a sparring partner
Suppose you have three propositions for a campaign. You can ask which feels relevant, what sounds unbelievable, what is confusing and which associations a phrase triggers. Then you can keep going. Ask whether the response changes in a different situation. Challenge the answer. Introduce another idea.
Traditional qualitative research is scarce. You recruit people, schedule sessions and have limited time with them. The client is watching and suddenly has another seventeen questions. Some can be squeezed in. A synthetic audience does not need to catch the train home.
That changes when research can happen. A half-formed hypothesis can be challenged before it becomes a recommendation. A creative team can explore whether an idea depends on an unexamined assumption. A strategist can test explanations for a behaviour before deciding which one deserves real research.
Real people are imperfect evidence too
Discussions of synthetic research often compare an imperfect simulation with perfectly truthful humans. People misremember, rationalise, protect their self-image and do not always know why they behaved as they did. An answer about why somebody bought something is information, although it cannot record the mental process that caused the purchase.
Synthetic respondents have no personal identity to protect, which can change the kind of answer produced. That is interesting, but it does not make the response equivalent to evidence from people who can actually buy, hesitate, misremember or surprise us. A model has no private truth to reveal.
Plausible is not the same as true
LLMs produce plausible explanations very well. That is why synthetic research can produce such useful material and why it can become dangerous.
Imagine a synthetic audience says people reject a proposition because it makes the brand feel too expensive. This is useful. It gives you something to investigate. Perhaps you ask follow-up questions, see the concern repeatedly and connect it with existing research.
The temptation is to put “consumers perceive the proposition as too expensive” into the presentation. We do not know that. We know the synthetic audience produced that response. Surprising answers need particular care because those are the ones people most want to promote into insights.
Evidence should determine confidence
There is no single rule for every decision. If I have good audience research, behavioural data and category knowledge, I am more comfortable using a synthetic audience to explore within that world. If I have almost nothing beyond “German men aged 25 to 34 who like sport,” I would be careful.
Commercial importance matters too. I might use synthetic responses to kill a weak headline before anyone spends more time on it. I would want stronger evidence before making a major investment because a simulation liked something. The confidence of the decision should reflect the quality of the evidence behind it.
Validation has to be empirical
We will understand the useful boundary through repeated comparison: synthetic responses against known human datasets, observed behaviour, research results and commercial outcomes. Validation needs to be specific to the category, population and task. A method that predicts familiar reactions to mainstream products may fail badly in a niche subculture.
Unknown unknowns are a serious limitation. If the relevant behaviour, language or community is poorly represented in training data, the model may produce a polished version of what outsiders already assume. The response can be coherent precisely because the missing reality never enters the simulation.
It can improve real research
One of the best roles for synthetic research may be before traditional research. Use it to explore the territory, find obvious answers and contradictions, discover which questions produce interesting responses and identify weak assumptions. Then spend scarce access to real people on the things you genuinely need real people to tell you.
I do not know where the boundary will settle as models, grounding and validation improve. For now, I use a simple rule. I am comfortable using synthetic audiences to explore. Claims about how people actually think or behave need evidence that justifies them. There is plenty of useful work between those points, and the synthetic audience really does have time for one more question.
