Research

A forecast is a claim. Advice is a claim. We grade both.

Live · updated daily

Forecast analysis

Every forecast we issue, graded against what happened.

We keep a permanent record of every forecast and score it against NOAA analysis after the fact — per model, per variable, per lead time. Where we do worse than guessing the local average, the number is negative and we publish it anyway.

Public Smarter Weather forecast skill scorecard
The public scorecard

In progress

Weather Bench

How well weather agents make decisions for real operations.

Can a weather agent do a staff meteorologist's job? We are building a public benchmark to find out: our agent is replayed against real past weather and graded on what a meteorologist on staff is paid for — speak up at the right lead time, stay quiet otherwise, reach the right people, cite the official source, and put honest odds on the call.

  • Real weather — past events rebuilt from public NOAA archives, never invented.
  • Sealed runs — the agent sees only what was knowable at the time. No web, no live data.
  • Honest grading — fixed rules, no model judging another model, confidence intervals, and the newest events held back. A human meteorologist baseline is planned.

Results publish when the first full run is scored.