Build Scope · v0.4 Greenfield 33 days to polls Hardened P0 shipped Plan committed

2026 Midterm
Forecast Engine

A daily probabilistic forecast for all 506 races on the 2026 ballot, and a dashboard that starts at “who controls the House” and drills to “which three polls moved NE‑02, and what does the majority look like if it flips.”

The model today · 2026-10-01

House 76.4% D (median 226 seats), Senate 61.2% D under the sit-out rule, governors median 21 D; 140 of 506 races carry a usable poll. This box and the masthead count are rewritten by every run (engine/prose.py). Everything else on this page is the scope as drafted on 22 August 2026 and its dated record of what shipped — the findings log in README carries everything since, and the dashboard’s “How it works” tab draws the model as it now stands.

Nov 3Election Day 2026
435House seats
35Senate seats · 2 special
36Governorships
218House majority
D+4Senate flip needed
The Board Verified against live sources, Aug 2026. Re‑verify before build.

What is actually on the 2026 ballot

Three simultaneous forecasts over one shared national environment. They differ in polling density, correlation structure, and how much of the answer the priors have to carry.

ChamberRacesControl mathPolling densityModeling character
House435 218 to win Very sparse — most districts have zero public polls Priors‑dominated. National environment × district lean does nearly all the work.
Senate35 D +4 for 51 Dense in ~10 battlegrounds, thin elsewhere Closest to a classic state model. Correlated error dominates the topline.
Governor36 No control stake Moderate, uneven Candidate‑driven, weakest partisan correlation, fattest idiosyncratic tails.

The one structural fact that shapes the whole House model

After the 2025–26 mid‑decade redistricting round, the median district is Virginia’s 1st, which Trump carried by 4.9 points — roughly 3.4 points to the right of the nation. Democrats can win the national House vote outright and still lose the chamber. Your seats‑votes curve has to encode that bias explicitly; a model that treats national margin as if it maps symmetrically onto seats will be systematically wrong in one direction.

Sanity anchor for the build

Public forecasts in early August 2026 put Democrats near 85% to win the House with a median around 230 seats (range ~211–253), and the Senate close to a coin flip. Treat these as a smoke test, not a target: when your v0 first produces a number, being wildly outside this range means the district priors or the sign conventions are wrong, not that you have found an edge.

Core Decision Make this call in P0. Everything downstream depends on it.

The simulation matrix is the product

Persist every draw, not the summary statistics. This is the single decision that separates a forecast you can drill into from one you can only look at.

Each run produces one array of shape n_sims × n_races holding a margin and a winner per cell. Every headline probability, every seat histogram, every tipping‑point statistic, and every conditional question is then a reduction over that array rather than a separate model:

# topline
P(House) = mean(seats[:, HOUSE].sum(axis=1) >= 218)

# conditional — the feature people actually come back for
mask     = win[:, "GA-SEN"] == DEM
P(Senate | GA-SEN goes D) = mean(seats[mask, SENATE].sum(1) >= 51)

# tipping point — sort each sim's races by margin, find the seat that crosses 218
# correlation view — corr(margin[:, i], margin[:, j])

Homegrown forecasters routinely store only per‑race win probabilities, then discover they can never answer a conditional question without re‑running the model. The storage cost of avoiding that is trivial:

ComponentShapeDtypeSize / run
Margins50,000 × 506float32~101 MB
Winners50,000 × 506int8~25 MB
Compressed (Parquet + zstd)——~25–40 MB
Full cycle, one run/day to Nov 373 runs—under 3 GB

One Parquet file per run date, partitioned by chamber. Precompute the L1/L2 aggregates into static JSON at run time; only conditional queries ever touch the matrix.

Architecture Five stages, strictly ordered. Each writes an artifact the next one reads — no stage calls another directly.

Engine pipeline

A batch pipeline, not a service. It runs once a day, writes immutable run artifacts, and the API only ever reads them. That makes every historical forecast reproducible and every regression bisectable.

01Ingest → raw_polls.parquet Pluggable source adapters. Append‑only, never mutate a fetched poll; corrections arrive as new rows.
02Normalize → polls.parquet Resolve race identity, dedupe across aggregators, tag population (LV/RV/A) and mode, reconcile candidate names.
03Estimate → posteriors.parquet National environment layer, then a per‑race mean and sd blending prior, polls, and fundamentals.
04Simulate → matrix.parquet Correlated Monte Carlo. Draws hierarchical error terms, tabulates seats and control.
05Serve → summary.json
→ races.json
Reduce the matrix into static aggregates; expose a thin conditional‑query endpoint over the raw draws.

The race registry is a first‑class input, not a constant

Districts, candidates, incumbency and map versions all change during the cycle — Texas’s map is still on appeal at the Supreme Court as of this writing. Keep districts_2026.csv and races.yaml versioned inputs with an effective date, so a late map change is a data update and a re‑run, not a code change. Never hardcode a district list.

Key every race on (cycle, map_version, district), never on the district string alone. Texas’s 35th district under the 2021 map and under the 2025 map are different electorates wearing the same name. A join on the bare string will silently blend them, and the resulting prior will look entirely reasonable.

Model · v0 Deliberately naive. Builds in days, not weeks, and forces the whole pipeline to exist before any of it is tuned.

The naive baseline, specified

This is what Phase 0 and 1 build. It is wrong in known, enumerated ways — and it produces a plausible, end‑to‑end forecast that the dashboard can be built against immediately.

### 1. Candidate roster  ────────────────────────────────
FEC bulk cn26 → party, office, district, incumbency (CAND_ICI)
# authoritative for federal races. Governors are not federal.

### 2. Resolve each poll to a matchup  ──────────────────
answers name PEOPLE, not parties: {"Casey Askar": 50.0}
  → surname match inside this race's roster, fuzzy at 0.85
  → drop matchups >45d behind the race's newest poll
    # TX-SEN polled Talarico-v-Paxton AND Talarico-v-Cornyn;
    # those are different races. Averaging them describes neither.

### 3. Weighted matchup average  ────────────────────────
w_p   = exp(-days_old / 30) · 0.35^partisan · 0.35^internal
eff_n = (Σw)² / Σw²                        # NOT the raw poll count
p_i   = Σ(w_p · margin_p) / Σ(w_p)

### 4. Prior — the fallback, not the driver  ────────────
lean_i   = pres24_margin_i - pres24_national
prior_i  = lean_i + generic_ballot + inc_i   # inc from FEC

### 5. Blend on effective n  ────────────────────────────
w_i   = eff_n / (eff_n + 3)
mu_i  = w_i·p_i + (1-w_i)·prior_i

### 6. Sigma conditional on PRIOR QUALITY  ──────────────
sigma_i = (1-w_i)·sigma_prior_only[chamber] + w_i·2.5
sigma_prior_only = {house 4.0, senate 6.5, governor 7.0}
    # a presidential-lean prior for a Senate race carries far
    # more error than a candidate-poll-informed one

### 7. Simulate, S = 50,000  ────────────────────────────
margin[s,i] = mu_i + 3.0·z[s] + 1.5·z[region,s] + sigma_i·t[s,i]
Two choices worth keeping from day one

Fat tails. Draw the idiosyncratic term from a Student‑t with df ≈ 5, not a normal. It is a one‑line change and it is the difference between a model that admits surprises and one that calls a 6‑point lead a certainty.

A hierarchical error, not independent races. Even the crude three‑level structure above captures the thing that actually breaks forecasts: polling misses are correlated. Independent race errors would give you a House probability of 99.9% and it would be nonsense.

What v0 knowingly gets wrong

  • No house effects. A pollster that leans three points red is counted at face value. Biggest single accuracy loss.
  • No pollster quality weighting. A partisan IVR poll counts the same as a gold‑standard live‑caller poll.
  • No fundamentals. Presidential approval and the midterm penalty — the most reliable signals in a midterm — are absent, so sparse‑poll races float on priors alone.
  • Elasticity fixed at 1.0. Real districts vary in how much they swing with the nation; college‑educated suburbs move more than ancestral‑partisan rural seats.
  • Crude correlation. A three‑level hierarchy, not a covariance matrix built from demographic similarity. Regions are a poor proxy for “races that miss together.”
  • Sign‑of‑margin decides the winner, which is simply false in Louisiana, in California and Washington top‑two generals that can be D vs D, and under ranked‑choice voting in Maine and Alaska — all of which have races in 2026.
  • No candidate quality, fundraising, scandal, or retirement effects.
  • Undecideds are allocated proportionally by silent default, which is what working in margins rather than vote shares amounts to. Fine in a clean two‑way race, wrong wherever undecideds run high or a third party is real. Make the rule explicit and configurable rather than implicit in the arithmetic.

Each of these is a named upgrade with a phase attached below. Shipping them in that order is deliberate: the accuracy gains are roughly monotonic down the list.

Upgrade Path Ordered by accuracy gained per day of work.

From baseline to a real model

CapabilityWhat it buysPhaseEffort
House effectsRemoves each pollster’s systematic lean by solving jointly for pollster bias and race means. Largest single win.P23–4 d
Own pollster ratingsWeight by historical error, computed from the archived poll corpus with shrinkage toward the mean. Removes the dependence on a third party’s ratings.P23 d
Fundamentals blendPresidential approval + midterm penalty + seat exposure, blended with poll weight rising as election day nears. Anchors the ~370 districts with no polls.P24 d
Covariance matrixCorrelated errors drawn from demographic distance (education, urbanicity, race, prior swing) instead of coarse regions. This is where topline calibration lives.P25 d
District elasticityPer‑district swing sensitivity estimated from prior cycles.P21–2 d
Special voting systemsRCV (ME, AK), top‑two generals (CA, WA), Louisiana. Correctness, not accuracy — the model is currently capable of reporting an impossible winner.P33 d
Candidate qualityPrior office held, fundraising, scandal flags. Real but modest; mostly matters in open seats.P32 d
Data The schedule risk lives here, not in the math.

Sources, and the hole where FiveThirtyEight used to be

Read this before estimating the project

FiveThirtyEight was shut down in March 2025. There is no longer a free, clean, actively maintained polls feed with stable schemas. Every homegrown forecast built since then spends more time on ingestion than on modeling, and plans that assume otherwise slip. Budget for adapter work and a manual poll‑entry path from the start — you will need it for under‑covered House races no aggregator tracks.

SourceProvidesAccessRisk
The Downballot
pres-by-CD, 2026 maps
2024 presidential result for every district under the new maps — the district priorFree, manual CSVCritical path No substitute exists
VoteHubLive race‑level and generic‑ballot pollsREST APIVerify Returned 403 to an unauthenticated probe — confirm auth, limits, and terms first
Silver BulletinGeneric‑ballot average, pollster ratingsPartly paywalledCite, don’t scrape
uspollingdata.comMeta‑average across six aggregatorsFree, dailyCross‑check only Never an input — it would launder others’ models into yours
MEDSL / Harvard Dataverse
10.7910/DVN/IG0UN2
House and Senate returns, 1976–2022Free .tab, GitHub mirrorStable The backtesting spine
538 GitHub archiveHistorical polls + legacy pollster ratingsFree, frozen at Mar 2025Stable Fine for backtests, useless for 2026
Wikipedia race tablesCoverage backstop for thinly‑polled racesScrapeMessy Inconsistent HTML, needs review queue
Ballotpedia / state SoSCandidates, filing, incumbency, retirementsSemi‑manualChurns Especially retirements
Cook / Sabato / DDHQRace ratingsManual entrySanity check Display alongside, never blend in

Licensing is a merge requirement, not a footnote

Every adapter records the terms it operates under, checked before it is written rather than after it is running. Wikipedia is CC BY‑SA and carries attribution obligations that propagate into anything public built on it; aggregator terms vary and some prohibit reuse outright. This costs an hour per source now and can invalidate a published product later.

Normalized poll schema

poll_id  source  pollster  sponsor  race_id  start_date  end_date
population {LV|RV|A}  mode {live|IVR|online|mixed}  sample_n  moe
dem_pct  gop_pct  other_pct  undecided_pct  margin
partisan_sponsor {none|D|R}  url  fetched_at  superseded_by

Everything else — weights, house‑effect adjustments, decay — is computed downstream and never written back onto the poll row. Ingestion stays dumb and auditable.

Dashboard Three levels. Every level answers a question the level above provokes.

Drill‑down structure

The information design goal: a visitor should be able to get a number in two seconds, understand where it came from in thirty, and interrogate it for an hour.

L1 — Overview

Wireframe · Landing
House · 85% Dcontrol dial
Senate · 52% Dcontrol dial
Governors · 21 Dmedian count
Democratic House seats — distribution of 50,000 simulations median 230 · 80% range 214–247
218 · majority 206 230 258 median 230

Sketch of the hero view, drawn on the early‑August consensus. Party color is reinforced by position relative to the threshold line, so identity never rests on color alone.

National environmentgeneric ballot trend · sparkline · approval
What changed todayprobability deltas, new polls, biggest movers

L2 — Chamber

Wireframe · House view
Hex cartogram — 435 districtsequal-area tiles, one per district
shaded by win probability
geographic area ≠ importance, so a
true map would be actively misleading
Path to the majorityraces sorted by forecast margin
cumulative seat count
tipping-point race marked
“the 218th seat is IA-03”
Race tablewin prob · forecast margin · rating · 7-day movement · poll count · last poll date — sortable, filterable, links to L3

L3 — Single race

Wireframe · Race detail
Forecast margin over timefan chart with 50/80/95% bands · individual poll dots sized by sample, opacity by recency · hover for pollster
Polls in this racepollster · rating · raw margin · house-effect adjustment · final weight — the adjustment shown, not hidden
Where the number comes fromprior / fundamentals / polls decomposition — how much of this forecast is actually data
Importancehow often this race is the tipping point across sims
Conditional“if this flips, control probability becomes X”

Cross‑cutting: scenario explorer and diagnostics

  • Scenario explorer. Pin outcomes for any subset of races; every other probability recomputes by filtering the simulation matrix. Effectively free once the matrix is persisted, and the single most‑used feature on forecasts that have it.
  • Correlation view. Which races move together, read straight off the matrix — makes the model’s error structure legible instead of a black box.
  • Model diagnostics page. Calibration curve, house‑effect table, backtest results, methodology writeup. If this is ever public, this page is what makes it defensible.
Interface Static where possible. Compute only for conditionals.

API contract

GET  /runs                          → available run dates
GET  /runs/{date}/summary           → topline probs, seat medians + intervals
GET  /runs/{date}/races             → per-race table (L2)
GET  /runs/{date}/races/{race_id}   → detail: history, polls, decomposition (L3)
GET  /runs/{date}/trend             → daily topline history for the fan charts
GET  /runs/{date}/matrix.parquet    → raw draws, for offline analysis

POST /runs/{date}/conditional
     {"given": {"GA-SEN": "D", "NC-SEN": "R"}, "ask": ["senate_control"]}
     → boolean-mask the matrix, reduce, return

The first five are files on disk, generated at run time and cacheable indefinitely — a run is immutable once written. Only /conditional needs a live process, and it is a masked mean over a memory‑mapped array: single‑digit milliseconds.

Quality Gate Not optional, and not last. It is how the variance parameters get set.

Backtesting and calibration

The σ values in the v0 spec are guesses. The only honest way to set them is to replay past cycles with data frozen at each date and check whether the resulting intervals were the right width.

Replay 2018 and 2022 — both midterms, both with a president of a known party, one a wave and one emphatically not. Freeze the poll corpus at each historical date, run the full pipeline, and score:

MetricQuestion it answersPass condition
Calibration curveDo 70% favorites actually win about 70% of the time?Within sampling error across buckets
Interval coverageDo 80% margin intervals contain the result 80% of the time?78–82% empirical
Seat‑count MAEHow far off is the median seat forecast?Beats uniform‑swing baseline
Brier scoreOverall probabilistic accuracy, per raceBeats uniform‑swing baseline

The baseline to beat is deliberately unglamorous: uniform national swing applied to the previous result in each seat. It is surprisingly hard to beat, and a model that cannot beat it is adding complexity rather than information. Nothing ships past Phase 2 until it does.

Watch specifically for over‑confidence at the topline: if coverage looks fine per race but the seat‑count distribution is too narrow, σ_nat is too small. That failure mode is exactly what produced the famous forecast misses, and it is invisible unless you score the aggregate rather than the races.

Score the competitive races, not all 435

A Brier score computed over every House seat is dominated by the roughly 370 districts that any model calls correctly at 99%. The engine and the baseline both score near‑perfect there, and the aggregate cannot tell them apart — a model could be badly wrong in every genuinely contested race and still post a fine overall number. Gate on the subset where the baseline gives between 5% and 95%. Report the full score too, but do not decide anything with it.

Two cycles cannot estimate a variance

This is the hardest limit on the whole plan and it deserves stating plainly. σ_nat is the parameter that most determines the topline, and replaying 2018 and 2022 supplies exactly two draws of the national error to fit it against. Fitting a variance to two observations is not estimation. So: take σ_nat from published polling‑error research as an informative prior, use the backtest to check that it is not too narrow, and never to shrink it. Add 2010 and 2014 for two more midterm draws if the data assembly is affordable. Where the backtest and the prior disagree, widen.

Schedule Each phase ends on a demonstrable gate, not a date.

Phase plan

P0

Data spine and walking skeleton

Race registry for all 506 races with ids, incumbents and map version. Downballot pres‑by‑CD loaded as district priors. One poll adapter working. v0 simulation running. One deliberately ugly HTML page printing a House probability.

GATE  A number comes out end to end, from raw poll to rendered page.
P1

Naive model complete, dashboard L1 and L2

All three chambers simulating. Run artifacts persisted daily on a cron. React dashboard with control dials, seat histogram, hex cartogram, path‑to‑majority, and the sortable race table. Staleness alerting from day one — a source that quietly starts returning zero rows must page someone, not sit there while the model runs on last week’s polls. Backtest data assembly starts here, in parallel. The forecast runs in shadow throughout, so that whenever it is first shown it already has a trend line behind it.

GATE  Reproduction test, not a consensus test. Feed the engine the actual 2024 national House margin under the 2024 maps and confirm it reproduces the actual 2024 seat split; then swap in the post‑redistricting maps and confirm the median district lands near R+3.4. This validates the priors and the seats‑votes curve against reality, offline, referencing nobody else’s model.
P2

The real model

House effects, pollster ratings, fundamentals blend, demographic covariance matrix, district elasticity. Backtest harness on 2018 and 2022, and the variance parameters tuned on it rather than assumed.

GATE  Beats the uniform‑swing baseline on Brier score and seat MAE, with 80% coverage in the 78–82 band.
P3

Drill‑down depth

L3 race pages with fan charts and poll decomposition. Scenario explorer and conditional endpoint. Tipping‑point statistics. Correlation view. Model diagnostics page. Special voting systems: RCV, top‑two, Louisiana.

GATE  “If Democrats win the Maine Senate race, what happens to Senate control?” is answerable in the UI in under five seconds.
P4

Freeze and operate

No structural model changes after roughly 20 October. Polls in, run daily, monitor for ingestion failures, publish the methodology page. Late‑cycle model changes are how forecasts embarrass themselves — if a flaw is found in the last fortnight, document it rather than patch it live.

GATE  A missed daily run is detected and alerted, not discovered.
—

Election night — explicitly out of scope

A live results and race‑calling model is a separate build with its own data contracts, its own costs, and a precinct‑level swing model that shares almost nothing with this one. It is not a stretch goal on this plan. Decide by 1 October whether to buy a results feed, build one, or simply go dark on election night.

The effort assumption, stated so it can be wrong

The sequence above assumes one engineer at close to full time, already fluent in Python numerics and React, with no other commitments. At half time, or split across other work, it does not fit — and the honest response is to cut scope against the ladder below rather than triage in the last fortnight. The phase and workstream dates this assumption was originally attached to have been removed: they forecast one particular pace, and the first three workstreams beat them by enough that the calendar had become misleading rather than useful. What still binds is the ordering, the gates, and the freeze.

Cut ladder — drop in this order

1Governors — all 36 races — no control stake, weakest correlation structure, cleanest amputation
2Hex cartogram → sortable table only — the table carries the information; the map carries the pleasure
3Scenario explorer UI — keep the conditional endpoint, which is nearly free once the matrix exists; drop only the interface
4Candidate quality signals — real but modest, mostly confined to open seats
5District elasticity — reverts to a fixed 1.0, which is wrong but unbiased
6Special voting systems → label ME, AK, LA and top‑two races “not forecast” — an honest gap beats a confident impossibility
7Model diagnostics page — only if this stays internal. If it is ever published, this rung is not cuttable.
Irreducible core — below this line, do not ship a forecast

Race registry · district priors · one working poll source · house effects · correlated simulation · the backtest gate · L1 and L2 views.

House effects are promoted out of Phase 2 into the core deliberately. “Naive baseline first” is the right way to get the skeleton standing, but an unadjusted poll average is the single most likely thing to be visibly wrong to anyone who knows the polls — and it is the cheapest serious upgrade on the list. Build naive; do not show anyone naive.

If the gate has not passed by the freeze

Decide the rule now, in August, not under pressure in late October. If the backtest gate has not passed by 20 October, the forecast is either labelled experimental with prominent uncertainty, or it is not published at all. Both are defensible. Publishing a forecast that failed its own quality gate, quietly, is not.

Layout Engine, API, and web are separately runnable. The engine never imports from the API.

Repository shape

election_predictor/ ├─ engine/ │ ├─ ingest/ votehub.py wikipedia.py manual_csv.py base.py │ ├─ registry/ races.yaml incumbents.csv districts_2026.csv │ ├─ priors/ pres_by_cd.py elasticity.py fundamentals.py │ ├─ estimate/ generic_ballot.py house_effects.py posterior.py │ ├─ simulate/ covariance.py draw.py tabulate.py │ ├─ backtest/ harness.py metrics.py cycles/{2018,2022}/ │ └─ runs/ {YYYY-MM-DD}/{summary.json,races.json,matrix.parquet} ├─ api/ FastAPI — routes.py conditional.py └─ web/ Vite + React + TS └─ components/ Dial SeatHistogram HexMap SnakeChart FanChart RaceTable

Polars over pandas for the poll and matrix work — the group‑bys and the Parquet round‑trips are the hot path, and it matters more than it usually would at this data size because you will re‑run the backtest harness hundreds of times.

Open Decisions Answers change the plan. Everything else I can decide as I build.

What I need from you

  1. Who is building this, and at what capacity? Every date in the phase plan rests on one full‑time engineer fluent in the stack. This is the answer that most changes the plan, and it is the one I do not have.
  2. Public or internal? A published forecast needs the methodology page, the calibration disclosures, and a much higher bar on house‑effect defensibility. An internal analysis tool does not.
  3. Any budget for paid data? Free‑only is entirely viable but slower to build and gappier in under‑polled House races. A paid feed mostly buys back ingestion time, which is the scarcest thing on this schedule.
  4. Governors in v1, or defer? 36 races, weakest correlation structure, no control stake. It is the cheapest thing to cut if the calendar tightens, and cutting it early is better than cutting it in October.
  5. Election night: build, buy, or skip? Needs deciding by 1 October, because building it means starting in P3.
  6. How many simulations? 50,000 is the assumption above. 10,000 is enough for toplines and cuts run time by 5×; conditional queries on rare scenarios get noisy below about 40,000.
Build Log P0 shipped 22 Aug. What the code taught us that the plan could not.

What shipped, and what it corrected

The pipeline runs end to end on real data. Five of the plan’s assumptions survived contact; five did not, and four of the failures were only visible once real polls were flowing.

The reproduction gate worked, and validated independently

The engine measures the median district at −3.10 — the published figure for the pre‑redistricting maps is Trump +3.1 — and 205 of 435 districts carried by Harris, the widely reported count. Neither was fed in. Replacing the consensus gate with a mechanical one (Stress Test #02) paid for itself on the first run.

Candidate-level, not generic ballot

The forecast is now driven by actual candidate matchup polling, with the generic ballot demoted to a fallback prior for the ~465 races that have none. Candidates come from the FEC bulk candidate master, which is authoritative for federal races. The effect was not marginal:

MeasureGeneric-ballot modelCandidate-level
Senate control13.9% D · range 49–5151.2% D · range 48–53
House control87.5% D · median 23386.9% D · median 229
Senate poll coverage0 of 35 races17 of 35 races
House incumbencyall UNK208 D / 211 R / 16 open

Five findings the plan did not anticipate

FindingWhat it brokeFix
A 14‑day half‑life discards nearly all the dataOH‑SEN‑SP had an effective n of 2.11 across 15 polls. The state forecast rested on two August releases; an NYT/Siena poll at −3 carried weight 0.0197. A state Trump won by 11 came out at 78.8% D.30‑day half‑life; shrinkage now uses effective n, never raw poll count. Ohio → 42.2%.
Distinct matchups were averaged togetherTX‑SEN carried 25 polls of Talarico‑v‑Paxton (+1.6 D) and 8 of Talarico‑v‑Cornyn (−1.0 D) — different races contingent on a primary. Their average describes neither.Drop matchups >45d behind the race’s newest poll. Fired on 10 races. TX‑SEN 57.7% → 42.2%.
The partisan flag was in the data and ignoredChange Research polls flagged PARTISAN:DEM counted at full weight.0.35 multiplier on partisan and internal polls.
FEC CAND_ICI means “holds the seat”It validated Senate incumbent party 34/35, but called all 8 flagged retirements “incumbent” — a retiring member still files as I.FEC overrides incumbent party only, never who is on the ballot.
Governors are not federalAll 226 governor polls resolve to no_roster; the FEC has no state coverage. Governor forecasts are prior‑only.Needs a non‑federal candidate source, or governors take cut‑ladder rung 1.

Two plan assumptions that were simply wrong

  • “The 538 archive is gone” was itself wrong — corrected. The repo is live (last pushed Feb 2025); the 404s were bad paths, not a deleted repo. pollster-ratings/raw_polls.csv holds 7,963 congressional general‑election polls with actual results attached. That single file unlocks house effects, our own pollster ratings, the empirical variance curve, and the backtest spine. Recording the correction because the wrong version briefly shaped the plan.
  • The generic ballot feed is 53 days stale at source. VoteHub’s newest generic‑ballot poll is 2026‑06‑30 while its approval polls run to 2026‑08‑18. Verified not to be an adapter bug. The staleness alarm from Stress Test #05 fired on day one, which is the only reason it was caught.
Known overconfidence — since fixed

Sigma did not widen with time to election, and it does now. Both remaining core items shipped: the error curve is fitted on 11,359 polls across 25 cycles, and every poll is de‑biased by its pollster. NC‑SEN went from 95.9% to the high 80s on unchanged evidence, and the whole model from a flat sigma of 4.18 to a fitted 5–13 per race. What replaced it as the leveraged number is the generic ballot’s own +2.74 point Democratic bias, worth about 13 points of House control per point on the current run — a figure that moves with the forecast rather than a constant, and one the dashboard now recomputes from the sweep on every run instead of quoting.

Committed Plan House and Senate only — governors cut at rung 1. All seven workstreams closed.

Road to parity

Calibration before coverage. The model’s biggest defect was not missing signal — it was overconfidence in the signal it already had. All seven workstreams are now closed, and the ordering paid for itself twice: the backtest’s own headline finding turned out to be an artefact of its stand‑in prior, and one workstream’s scoped mechanism was rejected on measurement rather than argued about.

Where it stood on 24 August 2026

Per‑district calibration sd(z) 0.85–1.08 on three replayed elections, against 1.74–2.03 for the flat sigma the engine shipped with. Four things were tested and cut on evidence: fundraising, candidate quality, demographic covariance and inverse‑variance poll weighting. An end‑to‑end audit pass then closed nine findings and four more it turned up, three of them live defects. The forecast that day: House 67.1% D, Senate 21.4% D — the Senate having moved on a one‑seat holdover correction that is decisive at a threshold of 51.

The measurement that set the order — the original diagnosis, since confirmed

Across 7,963 congressional general‑election polls with known results, poll error at 50–75 days out is RMSE 8.69 — Senate 7.95, House 9.95. The engine’s total sigma on a polled race is sqrt(3.0² + 1.5² + 2.5²) = 4.18. Averaging polls cancels noise but not house effects, which are systematic, so true forecast sigma at 73 days should be roughly 6–8. The model is about twice as confident as the data supports. That was the entire reason NC‑SEN reported 98%. The diagnosis held: measured leave‑one‑cycle‑out, the flat sigma’s sd(z) was 1.74–2.09, and its 80% intervals contained the truth 50–59% of the time.

Days to electionnMAERMSE
0–71,0075.016.42
21–351,7876.318.09
50–758946.818.69
A

Recalibrate sigma from the error curve

Fit sigma as a function of days‑to‑election and chamber on the 7,963‑poll corpus, replacing three guessed constants with a measured curve.

DONE  NC‑SEN landed at 86.9% on unchanged evidence. Leave‑one‑cycle‑out sd(z) 1.78–2.09 → 0.96–1.00.
B

House effects and our own pollster ratings

31 pollsters have 60 or more congressional polls in the corpus, with signed bias spanning −5.55 to +6.61. Every one is currently counted at face value.

DONE  499 pollsters, true lean sd 3.10, shrunk by empirical Bayes. Ratings score accuracy against the RESULT — the first version scored agreement with the field, which was circular and rewarded herding.
E

Backtest harness

Third, not last: without it, everything below is unfalsifiable. Replay 2018 and 2022 from the corpus, which carries poll date, poll margin, and actual margin together.

DONE  2018, 2020 and 2022 replayed through the live engine with the calibration held out. Curve and coverage reported; outcomes taken from all 435 districts, not from the poll corpus.
C

Presidential‑by‑district under the 2026 maps  gated

Clears the 113 prior_stale districts. Largest known House error: roughly two points of median‑district margin. Depends on a source that may be manual or licence‑restricted.

GATE TAKEN  No public source carries 2024 presidential results under the 2026 lines. Nothing was guessed. C was PRICED instead on the 2022 redraw: a prior joined across a redraw scores Brier 0.102 against 0.036, so C is a large prize. Sigma widened 3.82× as the interim — both ends of that ratio fitted, after the audit found the stored multiplier had drifted away from its own base.
D

Fundamentals for the unpolled districts

350 of 435 House districts have no candidate polling; their forecast is a prior wearing a probability. Adds fundraising (FEC data already in hand), candidate quality, and district elasticity.

DONE, PARTLY BY REJECTION  Elasticity beats it and transfers across three cycles. Fundraising and candidate quality were CUT on measurement — neither transfers between cycles.
F

Demographic covariance matrix

Correlated error from demographic similarity rather than a three‑level regional hierarchy. Fixes topline calibration, not point estimates.

DONE; GATE NOT DECIDABLE  Seat coverage is one observation per cycle and there are three, so it is reported rather than claimed: RMS seat miss 11.9 against a model seat sd of 14.4. The scoped mechanism was REJECTED — demographics add ±0.01 once state is removed; the missing layer was the STATE (sd 3.39, unmodelled).
G

Ranked‑choice and top‑two handling

Correctness, not accuracy. AK‑SEN sits in the competitive band today and is decided by sign‑of‑margin, which is simply the wrong rule for it. Also ME, LA, CA, WA.

DONE, one named exception  RCV and runoffs turned out to be EQUIVALENT — the polls used in ME and AK are head‑to‑head finals, which is the RCV final round. Louisiana’s jungle primary had already ended. California’s eight same‑party generals are enumerated. WA‑05 is named and forecast anyway.

Then P3, the drill‑down dashboard, and the freeze on 20 October, both unchanged.

Stress Test Adversarial pass over v0.1, 22 Aug. Findings folded back into the sections above.

Where the first draft broke

Fourteen failure modes, ordered by damage. Three of them changed the plan structurally rather than cosmetically, and those are worth reading in full.

Structural · 01  The schedule had no slack for its own named top risk

v0.1 identified ingestion as the primary schedule risk, then scheduled Phase 2 at roughly eighteen days of enumerated effort inside a twenty‑one day window, with no contingency anywhere in the plan. A plan that names a risk and then budgets as though it will not fire is not a plan, it is a forecast of the happy path. Added: an explicit effort assumption that can be checked against reality, a cut ladder that fixes the drop order in advance, and a stated rule for what happens if the quality gate has not passed by the freeze.

Structural · 02  The Phase 1 gate contradicted the Risks section

v0.1 gated Phase 1 on landing “within a few points of published consensus,” then warned four sections later against consensus‑chasing. A gate written in terms of other people’s models trains exactly the behaviour the risk section forbids, however carefully it is hedged. Replaced with a mechanical reproduction test that references nobody: run the engine on the actual 2024 national margin under the 2024 maps and confirm it returns the actual 2024 seat split, then swap in the post‑redistricting maps and confirm the median district lands near R+3.4. It tests priors and the seats‑votes curve directly, it runs offline, and it can be executed the day the registry is loaded.

Structural · 03  Two midterms cannot estimate a variance

The parameter that most determines the topline is σ_nat, and the backtest plan offered exactly two draws of the national error — 2018 and 2022 — to fit it against. Fitting a variance to two observations is not estimation, and the failure is silent: the model looks calibrated per‑race while the seat distribution is far too narrow. Changed to a prior‑first approach: take σ_nat from published polling‑error research, use the backtest only to check it is not too narrow, and widen rather than shrink where the two disagree. Source that prior properly during the build — do not inherit a number from this document.

The other eleven

Failure modeWhy it bitesChange
Brier over all 435 seats~370 districts are called correctly by any model; both engine and baseline score near‑perfect there, so the aggregate cannot distinguish them. A model could be wrong in every contested race and still post a fine number.Gate on races where the baseline gives 5–95%. Report the full score, decide nothing with it.
Backtest scoped as a task, not a data projectReplaying 2018 and 2022 needs two more registries — the 2011 and post‑2020 maps, plus 2016 and 2020 presidential‑by‑district — and a reconstruction of what was known on each date given publication lag.Promoted to a parallel workstream beginning in P1, not a P2 task.
Race IDs collide across map versionsTX‑35 under the 2021 map is a different electorate from TX‑35 under the 2025 map. A join on the district string blends them and produces a prior that looks perfectly reasonable.Composite key (cycle, map_version, district). Never a bare district string.
Conditionals go noisy on rare scenariosConditioning on a 5% event leaves 2,500 of 50,000 draws; two rare conditions leave dozens, and the UI will report them to one decimal place regardless.Surface effective n and Monte Carlo error on every conditional; refuse to display below n = 500.
Ingestion fails silentlyA source that starts returning zero rows does not error. The model keeps running on stale polls and the numbers simply stop moving.Staleness alerting from P1, not P4: page if a chamber’s newest poll exceeds its own 90th‑percentile gap.
Undecided allocation was implicitWorking in margins silently allocates undecideds proportionally — fine in a clean two‑way race, wrong wherever undecideds run high or a third party is real.Made an explicit, configurable rule; third‑party share carried as a field.
No shadow periodA forecast first shown in October has no trend line, and jumps visibly as the model settles — which reads as instability rather than as learning.Runs in shadow from P1, so it opens with history behind it.
Scraping terms never checkedWikipedia is CC BY‑SA with attribution obligations that propagate into anything public; aggregator terms vary and some prohibit reuse outright.Per‑adapter licensing check as a merge requirement, recorded in the adapter file.
P3 ended three days before the freezeA backtest failure surfacing late in P3 had nowhere to go except into the frozen window.Freeze is a gate, not a date; the cut ladder absorbs overrun.
Held‑over Senate seats invited hardcoding47 Democratic‑caucusing seats minus 12 up leaves 35 held over — until an appointment, vacancy, or party switch moves it, silently, mid‑cycle.Read from the registry each run; assert the chamber sums to 100.
“Beats uniform swing” may be near‑unreachable in the HouseWith ~370 unpolled districts the engine is largely uniform swing plus incumbency, so the gate could fail on similarity rather than on quality.Baseline specified exactly, scored on the competitive subset, and expected gains stated as modest rather than assumed large.
Risks

Where this goes wrong

  • Ingestion, not mathematics, is the schedule risk. The model in this document is a week of work. Getting clean, deduplicated, correctly‑attributed polls for 506 races every day, from sources that did not exist to serve you, is the rest of it.
  • Redistricting is still moving. Texas’s 2025 map is before the Supreme Court on appeal. Keep the map a versioned, dated input so a late ruling triggers a data change and a re‑run — never a code change under deadline.
  • Most House districts have no polls at all. The forecast for roughly 370 seats is a prior wearing a probability. Say so in the interface: show the poll count on every race, and let the decomposition panel make the ratio of prior to data visible rather than flattering.
  • Correlated error is where forecasts actually fail. Under‑specifying σ_nat produces a topline that looks confident and precise and is neither. Tune it on backtests. Never on intuition.
  • Consensus‑chasing. If v0 disagrees with published forecasts, the overwhelmingly likely cause is a bug. But once the backtest gate is passed, resist tuning toward other people’s numbers — at that point you are fitting to their models, not to elections.