TWO-WAVE STUDY
A computational mini-public
The better argument, measured.
Delphi Pythia gathers a public to argue a question, then measures two things: which arguments people rate as persuasive, and which ones change minds.
The problem
What the existing tools miss
Surveys measure what people believe. They are poor at explaining why anyone changed their mind. Deliberative forums surface real arguments, but they do not scale: a citizens' assembly is slow, costly, and small. Online discussion produces a great deal of text and very little measurement.
The same gap runs under all three. What people say and rate is an elicited signal, and it tells you which arguments sound convincing when someone is asked to judge them. What changes a position is a caused effect. The two come apart, and most instruments only ever see the first.
The solution
Discover, then verify
Delphi Pythia runs the deliberation and the experiment in the same instrument. On any contested question it works in two passes. Wave one discovers the arguments a public is raising, the positions they cluster into, and which of those positions is gaining ground as people rate one another's points. Wave two freezes that argument map and runs a randomized experiment on it, which shows which individual arguments are doing the moving.
That is what makes real deliberation practical at a scale citizens' assemblies cannot reach, without losing rigour. Separating each argument from its author means status and authority count for less. The analysis gives trustworthy estimates of how opinion moved, and it brings out the rare arguments that cross divides and win people over. It is built as a consultative instrument, a way to listen to a public more carefully.
How it measures
How the measurement works
Delphi Pythia is a computational mini-public, built on a long tradition of structured public deliberation: the Athenian assembly, Madison's call to refine and enlarge public opinion, Habermas's ideal of the unforced force of the better argument, and the anonymous, iterative rounds of the Delphi method. That lineage is why the instrument works the way it does. People contribute their own arguments, then read and rate everyone else's.
The first wave is open deliberation. People write arguments in their own words, and the system clusters them into an argument map as they arrive. An adaptive selector (exploration sampling, a Thompson-sampling variant) chooses which argument you see next, concentrating on the ones whose standing is still unresolved, while a built-in floor keeps every argument in rotation so none gets buried early. That tells us what people rate as persuasive, which is not always what moves them.
The second wave takes the same argument map, now frozen, and shows people randomized combinations of arguments, the way a lab experiment would. Position is measured before and after, so the shift can be read directly instead of inferred from a rating. Significance comes from reshuffling the real data thousands of times, which assumes nothing about the shape of the distribution, and the ranking reports its own uncertainty.
The demo case
Where the debate is heading
The instrument is running end to end on a demonstration case: how AI should be governed. See the live results →
Live counts from the simulated demo — an AI panel.
These findings come from our first human pilot: 112 people worked through the question of how AI should be governed. About one in six (17%) changed their stated position along the way, and average confidence in the position people landed on rose by about half a point on the five-point scale. A few patterns are already clear.
Initial vs. final position, first human pilot · n = 112
Over the course of this pilot, support moved toward an international body to coordinate AI rules. Of the three positions on the table, it gained the most ground.
-
Accountability is what people rate highest.
The argument rated most persuasive across the divide was that companies should not be allowed to run AI without legally binding government oversight. Even people who arrived favouring a different approach rated it 3.9 / 5. The weakest was the call to let industry regulate itself.
-
There is real common ground.
Two arguments are rated well by every camp: that a single international body should set common rules all countries follow, and that companies should face clear standards, public transparency, independent audits, and real liability. Either could anchor a position most people could live with.
-
Rated highest is not the same as most effective.
This pilot measured what most polling and message-testing measures: which arguments people rate highly after reading them. It does not tell us which of those arguments moved someone's position, as opposed to simply being on screen when the position moved. That is what the second wave is for: a randomized experiment on the same argument map. See it running now, above.
About half of the AI panel, –, changed position between first and final answer, and confidence in the position they landed on rose by about – points on an 11-point scale.
Initial vs. final position, simulated demo · AI panel · n = – · live
-
Rated most persuasive right now.
– Rated – on average across the panel.
-
Where the camps agree.
– holds the highest floor across all three camps — even its lowest rating from any side is –.
-
Rated highest is not the same as most effective.
The randomized second wave has so far verified – of – frozen arguments as moving positions; the strongest shifts stances by – points on the movement scale.
Before this simulated run, a first human pilot reached the same headline: among 112 people, support for an international body rose from 45% to 53%, about one in six switched position, and the top argument was rated 3.9/5 even by other camps. The run above is the same instrument, run at speed.
The same two-wave process can run on any contested question, from a policy to a product to a referendum: discover the arguments a public is making, see which way opinion is moving, and verify which specific arguments are shifting positions. Ratings alone cannot do the last part.
The cases
The same instrument, different questions
Each one runs the same two waves on a different contested question. Pick one to see how its argument map moved, or take the study yourself.
Take one of the studies yourself. You state a position, weigh the strongest arguments from every side, and watch where you land as the room moves.
or see the live results 📊