Reasoning effort
| Seed 2.0 Lite | gpt-oss 120B | Codex 5.3 | |
|---|---|---|---|
| None | — | — | |
| Low | |||
| High | — | ||
| Default |
Traditional benchmarks ask: “Can you recall the answer?”
Forecasting = Gathering × Synthesis × Judgment × Decision
The core composite capability driving LLMs toward decision support
May every forecast be reproducible, may AI truly become decision support
In service of every person’s judgments and choices for a good life
Best score per model
| Score rank | Model | Reasoning | Score / 100 | Cost / 300 questions | Mean time / answer |
|---|---|---|---|---|---|
| 01 | High | 68.67 64.44–72.74 | $36.17 | 102s | |
| 02 | ON | 68.15 63.63–72.35 | $64.74 | 78s | |
| 03 | Default | 67.85 63.33–72.02 | $42.53 | 74s | |
| 04 | Default | 66.16 61.81–70.38 | $0.49 | 66s | |
| 05 | High | 65.72 61.16–70.03 | $48.75 | 146s | |
| 06 | Low | 64.73 60.14–68.90 | $19.26 | 41s | |
| 07 | Default | 64.65 60.18–68.94 | $1.36 | 89s | |
| 08 | High | 64.18 59.50–68.60 | $0.73 | 71s | |
| 09 | High | 63.55 59.01–67.75 | $1.99 | 110s | |
| 10 | Low | 62.58 57.91–67.01 | $1.19 | 71s | |
| 11 | None | 62.56 57.99–66.88 | $1.45 | 77s | |
| 12 | Default | 62.36 58.01–66.72 | $21.72 | 62s | |
| 13 | OFF | 62.34 58.07–66.56 | $2.04 | 94s | |
| 14 | Minimal | 62.01 57.21–66.75 | $1.99 | 21s | |
| 15 | ON | 61.93 57.62–66.33 | $10.96 | 94s | |
| 16 | OFF | 61.62 57.42–65.59 | $8.89 | 92s | |
| 17 | ON | 61.45 57.25–65.66 | $0.55 | 77s | |
| 18 | ON | 60.91 56.62–65.25 | $1.02 | 151s | |
| 19 | ON | 60.72 56.10–65.07 | $15.38 | 69s | |
| 20 | ON | 60.68 56.38–64.80 | $1.75 | 186s | |
| 21 | ON | 59.89 55.35–64.34 | $14.87 | 95s | |
| 22 | ON | 59.27 54.66–63.50 | $4.46 | 82s | |
| 23 | Default | 58.62 54.28–62.94 | $4.11 | 87s | |
| 24 | OFF | 58.52 54.42–62.53 | $0.51 | 52s | |
| 25 | OFF | 58.17 53.79–62.52 | $0.93 | 91s | |
| 26 | Low | 57.70 53.43–62.02 | $0.22 | 37s | |
| 27 | Default | 57.67 52.86–62.15 | $8.43 | 48s |
No matching models.
Cost: estimated forecasting inference in USD, excluding search and auditing. Intervals: 95% CI.
| Seed 2.0 Lite | gpt-oss 120B | Codex 5.3 | |
|---|---|---|---|
| None | — | — | |
| Low | |||
| High | — | ||
| Default |
| OFF | ON | Δ Score | |
|---|---|---|---|
| Qwen3.5 Flash | +2.93 | ||
| Qwen3.5 Plus | -1.66 | ||
| Qwen3.5 35B | +2.74 | ||
| Kimi K2.5 | +0.31 |
SCORE / 100
68.67 362.59 / 528 pointsYes/no, binary, and single-answer questions earn full credit for a correct answer and zero otherwise. Multi-answer questions use F1 = 2TP / (2TP + FP + FN).
Invalid answers and refusals score zero. Only complete runs are ranked.
95% confidence intervals: 2,000 bootstrap samples, stratified by question type and clustered by question.
Search results are filtered by the declared knowledge cutoff and independently checked for leakage. Native model browsing is disabled.
Per answer: up to 6 reasoning steps, 4 searches, and 5 results per search.
Event resolution: 12 March–21 September 2026.
GPT 6.1 Sol · 137 / 10,000
Human review confirmed all 137 flagged leaks and found none among 1,000 results labelled non-leaking.
Dataset