What the market asked
The Polymarket event was straightforward in wording but narrow in scope: “Iran agrees to end enrichment of uranium by July 31?” Categorized under Politics, the market turned on a binary resolution. According to the stated rules, YES required Iran to publicly agree or pledge to end all uranium enrichment by the deadline of July 31, 2026, 11:59 PM ET. Anything short of that public, comprehensive commitment resolved NO. The council reviewed the market on 2026-07-25, six days before the cutoff.
What the council concluded and why
All four models—Gemini, Grok, Claude, and GPT—participated, with Grok acting as chairman. The council’s action was BUY_NO. It assigned a 1% probability to YES and expressed 92% confidence in that assessment. The published reasoning was explicit:
“Council recommends BUY_NO (NO). Probability: 1%, Confidence: 92%. Resolution criteria (per Polymarket rules): YES only if Iran publicly agrees/pledges to end all uranium enrichment by July 31, 2026 11:59 PM ET.”
The emphasis on the word “all” and the requirement for a public pledge framed the analysis. The models treated the bar as high: diplomatic language, partial freezes, or unconfirmed reports would not suffice. Given the short remaining window and the historical pattern of Iranian nuclear negotiations, the council saw virtually no path to the precise YES condition.
What the market itself was pricing
The graded report does not record the prevailing Polymarket mid-price or order-book depth at the moment of the council’s decision. What is recorded is the council’s own probability: 1% for YES. That figure represented an extreme view that the contract was, for practical purposes, already decided against the YES side under the written rules. The recommendation to buy NO flowed directly from that probability and the 92% confidence attached to it.
What actually happened
The market resolved NO. Iran did not issue a public agreement or pledge to end all uranium enrichment by the deadline. The council’s BUY_NO action was therefore correct. Graded on the action basis, the result produced a Brier score of 0.0001—essentially the minimum possible error for a near-certain forecast that matched the outcome. The low score confirms both calibration and the decisive role of the resolution criteria: the event did not hinge on broader geopolitical atmospherics but on whether a specific public commitment appeared before the timestamp.
Takeaway
This case illustrates the value of reading resolution language literally. The council’s 1% probability was not a geopolitical forecast of Iranian intentions in the abstract; it was a narrow judgment that the exact YES trigger was vanishingly unlikely inside the remaining calendar window. When markets use precise, public-action thresholds, models that stay tightly coupled to those thresholds can produce highly accurate calls even on contentious topics. The near-zero Brier score here is less a story of brilliant foresight than of disciplined adherence to the rules that actually paid out.
AI-generated analysis for informational purposes only. Not financial advice.