Traditional benchmarks ask: “Can you recall the answer?”

OracleProto asks: “Can you predict the future?”

Forecasting = Gathering × Synthesis × Judgment × Decision

The core composite capability driving LLMs toward decision support

May every forecast be reproducible, may AI truly become decision support

In service of every person’s judgments and choices for a good life

Model performance

Best score per model

Download results

Score · 0–100

95% confidence intervals

100 75 50 25 0
68.7 GPT-5.4 High
68.1 Opus 4.6 ON
67.9 Sonnet 4.6 Default
66.2 gpt-oss 120B Default
65.7 Gemini 3.1 Pro High
64.7 Codex 5.3 Low
64.7 Seed 2.0 Lite Default
62.3 Qwen3.5 Plus OFF
62.0 Gemini 3.1 Lite Minimal
61.9 Kimi K2.5 ON
61.4 Qwen3.5 Flash ON
60.9 Qwen3.5 35B ON
60.7 GLM-5 ON
59.3 MiniMax M2.5 ON
58.6 DeepSeek V3.2 Default

Full ranking

27 / 27
Score rank Model Reasoning Score / 100 Cost / 300 questions Mean time / answer
01 High 68.67 64.44–72.74 $36.17 102s
02 ON 68.15 63.63–72.35 $64.74 78s
03 Default 67.85 63.33–72.02 $42.53 74s
04 Default 66.16 61.81–70.38 $0.49 66s
05 High 65.72 61.16–70.03 $48.75 146s
06 Low 64.73 60.14–68.90 $19.26 41s
07 Default 64.65 60.18–68.94 $1.36 89s
08 High 64.18 59.50–68.60 $0.73 71s
09 High 63.55 59.01–67.75 $1.99 110s
10 Low 62.58 57.91–67.01 $1.19 71s
11 None 62.56 57.99–66.88 $1.45 77s
12 Default 62.36 58.01–66.72 $21.72 62s
13 OFF 62.34 58.07–66.56 $2.04 94s
14 Minimal 62.01 57.21–66.75 $1.99 21s
15 ON 61.93 57.62–66.33 $10.96 94s
16 OFF 61.62 57.42–65.59 $8.89 92s
17 ON 61.45 57.25–65.66 $0.55 77s
18 ON 60.91 56.62–65.25 $1.02 151s
19 ON 60.72 56.10–65.07 $15.38 69s
20 ON 60.68 56.38–64.80 $1.75 186s
21 ON 59.89 55.35–64.34 $14.87 95s
22 ON 59.27 54.66–63.50 $4.46 82s
23 Default 58.62 54.28–62.94 $4.11 87s
24 OFF 58.52 54.42–62.53 $0.51 52s
25 OFF 58.17 53.79–62.52 $0.93 91s
26 Low 57.70 53.43–62.02 $0.22 37s
27 Default 57.67 52.86–62.15 $8.43 48s

Cost: estimated forecasting inference in USD, excluding search and auditing. Intervals: 95% CI.

Reasoning and score

Reasoning effort

Seed 2.0 Lite gpt-oss 120B Codex 5.3
None — —
Low
High —
Default
Score 55 70

Thinking ON / OFF

OFF ON Δ Score
Qwen3.5 Flash +2.93
Qwen3.5 Plus -1.66
Qwen3.5 35B +2.74
Kimi K2.5 +0.31

Model details

GPT-5.4

High

SCORE / 100

68.67 362.59 / 528 points
Yes / no 100 questions · 100 pt
74.33
Named binary 9 questions · 9 pt
77.78
Single-answer choice 154 questions · 308 pt
65.15
Multi-answer choice · F1 37 questions · 111 pt
72.60
Multi-answer exact accuracy 38.74%
Exact answers 65.33%
Repeat agreement 72.22%
Same wrong set · all 3 12.00%
Sources gpt-5.4-high--provider-default

API pricing

Evaluation method

Evaluation protocol
300 Resolved questions
528 Available points
3 Repetitions / question
27 Model configurations

Scoring

Score = Mean points over 3 runs 528 × 100
Yes / no 1point / question 100 questions · 100 pt
Named binary 1point / question 9 questions · 9 pt
Single-answer choice 2points / question 154 questions · 308 pt
Multi-answer choice · F1 3points / question 37 questions · 111 pt

Yes/no, binary, and single-answer questions earn full credit for a correct answer and zero otherwise. Multi-answer questions use F1 = 2TP / (2TP + FP + FN).

Invalid answers and refusals score zero. Only complete runs are ranked.

95% confidence intervals: 2,000 bootstrap samples, stratified by question type and clustered by question.

Evaluation setup

Search results are filtered by the declared knowledge cutoff and independently checked for leakage. Native model browsing is disabled.

Per answer: up to 6 reasoning steps, 4 searches, and 5 results per search.

Event resolution: 12 March–21 September 2026.

Search-result leakage

1.37%

GPT 6.1 Sol · 137 / 10,000

Human review confirmed all 137 flagged leaks and found none among 1,000 results labelled non-leaking.

Dataset