Skip to main content

Research Lab

The Lab studies outcomes so the next forecast starts sharper.

Every resolved question is evidence about the method that produced it. The Lab turns that evidence into changes in how D weighs evidence, evaluates experts and identifies Signals.

  1. 01Evidence6,000+ sources, deduplicated and dated.
  2. 02Judgment1,000+ ranked experts, weighted by past accuracy.
  3. 03ForecastA price on every market D tracks.
  4. 04OutcomeResolution recorded against the forecast.
  5. 05Research LabOutcomes studied for what moved and what missed.
  6. 06FindingsWeights updated. The next forecast starts sharper.

What the Lab tests

Three questions, asked continuously.

Call accuracy and calibration

When D says 70%, does it happen 70% of the time? Every closed forecast is scored against its stated probability and binned into a calibration curve.

Signal outcomes and timing

Edge is only half of a Signal. The Lab measures whether the near-term reason actually moved the market inside the exit window, and retires the reasons that stop working.

Forecasting beyond markets

Not every question has a contract. The Lab tests whether the same method holds on questions no venue lists, so the system improves faster than the markets it reads.

Calibration

When D says 70%, does it happen 70% of the time?

Closed forecasts are binned by the probability D stated, then compared with how often those questions actually resolved yes. The bars below are stated against actual.

0–10%stated 5% · actual 4%
10–30%stated 20% · actual 22%
30–50%stated 40% · actual 37%
50–70%stated 60% · actual 63%
70–90%stated 80% · actual 78%
90–100%stated 95% · actual 96%

Featured report

Whose Iran calls held up?

0

Predictions scored

0

Experts ranked

0

Months of record

A retrospective scoring of every public Iran forecast made between January and July, ranked by Brier score and grouped by the kind of evidence each expert leaned on.

Read it in the app