Colorado River/course project
The same snowpack forecasts two reservoirs differently
Lake Powell and Lake Mead sit on the same river, fed by the same upstream snow. A model trained on 35 years of SNOTEL data beats every naive rule at one and loses to a seasonal rule at the other. On 17 April 2026 both forecasts broke at once.
From 2014 to 2024 my family lived in Henderson, Nevada, where roughly 90 percent of the water supply is Colorado River water drawn from Lake Mead. Southern Nevada cut per capita use 58 percent between 2002 and 2025 while adding approximately 876,000 residents. Lake Mead fell 173 feet over the same period. Nevada holds the smallest Colorado River apportionment of the three Lower Division states, so local conservation, however impressive, cannot move basin storage materially.
That gap between local effort and basin outcome is what made me want to forecast the reservoirs directly rather than reason about them.
The question
Snow falling in the Rocky Mountains is measured daily by the SNOTEL network and eventually becomes water in Lake Powell and then, after release through Glen Canyon Dam, water in Lake Mead. If snowpack drives storage, a model given 35 years of snow data should forecast pool elevation better than a rule that just extrapolates the reservoir’s own recent behaviour.
Three such rules are the bar to clear. Persistence says tomorrow’s elevation is today’s. Drift extends the last 30 days of trend forward. Seasonal naive adds the average change observed on the same calendar days in the training record. Beating a model is easy; beating the strongest of these three is the actual test.
The figure
Every cell below is a real out-of-sample forecast: XGBoost trained on elevation change over 1990–2022, feature set chosen on a 2023–2024 validation split, then run untouched on January 2025 through 6 August 2026.
Loading figure…
What it shows
At Powell the model wins everywhere. It beats the strongest naive rule at all three horizons, with skill scores of +0.78, +0.71, and +0.62. The seasonal rule fails badly here because the 2025–26 spring runoff never arrived. A rule built on the average year cannot represent a year that didn’t happen.
At Mead the model wins only at 30 days. Past that, seasonal naive is better: 3.16 ft against 3.46 at 60 days, 4.34 against 5.39 at 90. The reason is that Mead’s annual cycle is largely administrative rather than hydrological. Releases follow an operating schedule, and a rule that memorises the calendar is memorising the schedule.
The ablation is the interesting part. Training on SNOTEL and calendar features alone, with no autoregressive lags at all, beats the full model at Powell 60 and 90 days, and at Mead 90 days. Autoregressive features were calibrated in a wetter regime; in 2025–26 the drought signal in snowpack generalises better than momentum in the reservoir level. Validation nonetheless selected the full feature set in every cell, and that is what the headline numbers report. Choosing SNOTEL-only after seeing the test results would be choosing on the test set.
What no model caught
On 17 April 2026 the Department of the Interior cut Lake Powell releases from 7.48 to 6.0 million acre-feet. Toggle the figure to Lake Mead at 90 days and watch the lower panel: the mean residual moves from +1.12 ft before that date to −8.60 ft after it. The model was already forecasting slightly high; afterwards it was forecasting a reservoir that was no longer being filled on the old schedule.
Water held back at Powell is water that does not arrive at Mead, which is why the residual steps in opposite directions at the two reservoirs. This is not a tuning failure. Scheduled operations are learnable from history. A discretionary decision is not, and the shaded region marks the honest boundary of what this model can claim.
What I got wrong the first time
The April report claimed R² ≈ 0.98. Future target columns had been left in the feature matrix, so the model was reading a copy of its own answer. Two symptoms should have caught it before submission: test error came in below validation error on an out-of-distribution drought period, and error did not increase with forecast horizon. Both are structurally impossible for an honest forecast.
Correcting it took three changes made together:
- Current elevation restored to the feature matrix. Elevation at time t is not leakage when forecasting t+h. Excluding it while the baselines used it made the comparison unfair to the model.
- Delta target. Gradient boosting cannot extrapolate below its training
range, so the model now predicts the change
y(t+h) − L(t), which also aligns it with the delta baselines. - Valid-date axis. Residuals and figures are plotted on the date a forecast applies to, not the date it was issued.
The corrected 30-day Powell RMSE is 1.31 ft, against the retracted 1.36 ft. The numbers are comparable because a 30-day change anchored on today’s elevation is genuinely a slow problem for this reservoir, and the ablation shows where the remaining skill comes from.
Reproducing it
The pipeline runs in six steps from raw download to publication figures, and the site’s copy of the data is generated from the same artifacts by a single export script. Nothing on this page is transcribed by hand.
Materials
- Submitted report (April 2026)report
- Errata (August 2026)errata
- USDA NRCS SNOTELdata
- USBR HydroDatadata