A fundamental flaw of prediction markets
A while ago, I wrote an article in the Financial Times about the failures of prediction markets. In it, I said that one of the potential sources of failure is that prediction markets give more weight to people who bet more money on their forecasts. But who is to say that people with more money or higher conviction are necessarily more accurate?
Not to pat myself on the shoulder, but a new study by a group from UC Berkeley investigated the microstructure of Kalshi and Polymarket predictions and demonstrated that this is exactly the problem. People who bet more money in prediction markets (often called ’whales’) aren’t more accurate than the average or smaller participants. They are, in truth, worse, which makes the predictions of prediction markets flawed in practice.
The researchers collected the data on a total of 5,456 markets on Polymarket and Kalshi. They grouped them into 15-minute crypto markets (e.g., “Will Bitcoin be up or down in the next 15 minutes?”), Mention Markets (e.g., “What will Amazon say during their next earnings call?”), and forecasts for NBA games (I sense that some of the authors were betting on NBA games and needed an excuse to check if prediction markets help there).
It’s a very comprehensive paper that looks at the bias in prediction markets and finds that they do have a systematic bias, but that bias can sometimes be optimistic and sometimes pessimistic, depending on the market. But it also looks at the ‘edge’ traders have as a function of the size of their wagers. Edge, in this case, means how profitable their trades are on average. If ‘whales’ have a larger or at least the same edge as the average participant, then prediction markets should be a reliable forecasting tool. If their edge is smaller than the average trader’s edge, then the aggregate forecast of the prediction market is unreliable and worse than the ‘wisdom of crowds’ generated by an unweighted average of opinions.
The charts below show the average edge of the traders sorted by the size of their trades from smallest (left) to largest (right).
Average edge of traders by size of trade
Source: Daleep et al. (2026)
In 15-minute crypto markets, the smallest traders have the largest edge, while the largest traders have the lowest edge. In Mention Markets, the whales even have a negative edge, i.e. following their trades tends to lose you money. But because they are the ones that influence the aggregate the most, it means that following the prediction market forecast tends to lose you money or, at the very least, not make you money. In the NBA predictions, the picture is similar to mention markets where the largest traders have a negative edge and, on average, lose money to the smaller traders.
Further analysis indicates that the larger traders have the power of their conviction but not the power of information. Indeed, larger traders seem to act mostly on ideology rather than information, and this ideology informs their high conviction. Or, to quote the authors in their abstract: “We find no statistically significant correlation between sentiment intensity and informational edge, indicating that the most prominent voices in these ecosystems function primarily as sources of communicative noise.”



I have a question that might interest many people.
Let’s apply this line of reasoning to financial markets—specifically to E-mini S&P 500 options, for which the CME provides daily data. I’ve met people who have spent years analyzing block trades and the CME Globex Trade Browser (attached below). The goal is to try to understand what the "whales" are doing in the short term and, by looking at the strategies they’ve deployed, infer the market's direction for the day or the week.
While you can determine the direction (buy vs. sell) of block trades, Globex trades do not carry a directional indicator, meaning you have to speculate on their directionality.
A second approach involves analyzing open interest across various option expiration dates. Strike prices are plotted on the horizontal axis, and a histogram is drawn for each strike. The height of each bar is proportional to the number of open contracts at that strike, broken down by calls and puts. A static analysis reveals support and resistance levels, while a dynamic analysis suggests how the whales are behaving in response to movements in the underlying asset.
Are there any studies on whale activity acting as a leading indicator? And what would our hypothesis be? Would we see results similar to those of the Polymarket study, or do financial markets have their own behavior?
I hope I’ve made myself clear. It’s a topic that really interests me.
Thank you, Joachim
EDIT:
HERE ARE THE LINKS:
https://www.cmegroup.com/qa/block-trade-browser-old.html
https://www.cmegroup.com/reports/daily-index-option-spread-activity.pdf
https://www.cmegroup.com/tools-information/quikstrike/cme-globex-trade-browser.html
Fascinating — and the negative-edge-for-whales is striking.
Two things I'm chewing on. First, persistence: if the whale mispricing is this consistent, what stops it being arbitraged away? My guess is it's a limits-of-arbitrage story — the conviction-driven flow is persistent and betting against it carries real timing/resolution risk — but I'd love to know if the paper takes a view on why smart money doesn't simply fade the whales.
Second: within each market type, do markets with a heavier whale concentration underperform comparable markets with flatter participation — once you account for how inherently (un)predictable each event is? The across-type differences are clear; I'm curious whether whale-concentration itself is the variable, holding predictability constant.
The design implication almost writes itself — an unweighted, or size capped "wisdom of crowds" benchmark running alongside the money-weighted price — though as you note, that's essentially what the paper shows already beats the whales.
Care to build this, anyone? :)