← Back to the articleSnapshot through 2026-09-19 · All 70 station summaries · No individual forecast recordsAbout this edition
FORECAST RESEARCH / ALL-STATION ARCHIVE

A Nationwide Dive Into Forecasting Accuracy

Explore every available station, with comparable histories ranked over shared calendar months.

Station explorer ↗Get the dataset ↓

We began saving weather forecasts without being sure what we would find. Years of repeated forecast collection now let us follow how temperature predictions change from six days ahead to the near term, across stations around the United States and Puerto Rico. The clearest finding in this archive is geographic: middle America generally appears harder to predict than coastal America, with larger differences from the near-term forecast, especially several days ahead. We compare the same sampled months across stations and retain an audit of flagged readings. This pattern describes our archived forecasts; it does not establish accuracy against observed weather or explain what causes the regional differences.

Agreement with the near-term forecast. These comparisons do not establish observed-weather accuracy. An unusual station can reflect geography or forecasting difficulty; it does not establish bad data.

Export all rankings ↓

Loading the station analysis…

GEOGRAPHY / STATE COMPARISON

Where forecasts differ most.

Mean absolute difference from the near-term forecast.

Loading state boundaries…

Mean absolute difference · fixed scale across years, leads and quality views
No comparable data

The map uses available pairs at the selected lead. Choose matched year to date to hold stations and calendar coverage fixed across years. A year-to-year change alone does not establish improving forecasts. Rankings and charts below retain their completed-year window and require all six leads at the same target hour.

Each contributing station receives equal weight within its state, after equal weighting of the selected calendar months. Thin coverage means fewer than 30 pairs or 10 sampled dates in a month. These are summaries of sampled stations, not statewide accuracy estimates. Alaska, Hawaii and Puerto Rico appear as insets. Boundaries: U.S. Census via us-atlas ↗

Download map summaries ↓

State values and coverage
01 / OUTLIER SCREEN

The stations to investigate.

Potential outliers use |modified Z| > 3.5; the broader watchlist uses |ordinary Z| > 2.

See flags across all six leads and both quality views

Flags consider five correlated metrics at the selected lead. They are exploratory labels, not adjusted significance tests. A station stays in the analysis when flagged.

02 / COHORT RANKING

Mean absolute difference

Equal weights for the same sampled year-months. Select a bar to inspect a station.

03 / MAGNITUDE & DIRECTION

Different ways to disagree.

Horizontal: magnitude. Vertical: signed bias. Outlined points are flagged on any metric.

Positive bias = warmer than the zero-hour reference. Hover, focus the table below, or select a point for details.

04 / INSPECT A STATION

Station details

Explore station summaries ↗
05 / ALL THE EVIDENCE

Comparable stations.

Select a station name to inspect it. Every column reflects the selected lead and quality view.

06 / ARCHIVE COVERAGE

Every station has a history.

Export coverage ↓

View all station histories and comparison eligibility
Selection, weighting & interpretation

Archive and comparison eligibility

Include all station records with forecast history, except the explicitly identified duplicate Salt Lake City #80. Keep Salt Lake City #1. Other same-city records remain distinct and display their IDs. Short or incomplete histories remain downloadable and explorable.

A fairer calendar comparison

Use the latest three completed calendar years. Each retained site/month/lead must have at least 30 screened pairs on 10 local dates. Use the same supported year-months at every station and lead; average each month's metric with equal weight. Each station's target hours must have valid pairs at all six leads. Exact hours still differ across stations.

Reading-level and station-level flags

The existing two-sigma temperature screen is unchanged. Flagged target-hour share means any of the six leads or its reference was flagged. Station-level scores compare these summary metrics across the cohort; no station is deleted.

Uncertainty and limits

Pointwise 95% intervals use 1,000 paired year-month block resamples within calendar-month strata. The seasonal mix stays fixed. Peer-gap intervals compare a station with the other comparable stations' median. Neighboring months may still be dependent, and intervals do not correct sampling bias, regional differences or multiple comparisons.

The modified Z score uses the median and median absolute deviation. NIST outlier guidance ↗

Download written analysis ↓ · Full analytical results ↓