Baruch College Pre-MFE · Machine Learning · 2026
CPI inflation forecasting
Can macro data predict next month's US inflation? OLS, Ridge and Lasso trained on 28 engineered FRED features from 2008 to 2022, tested on 2023 to 2025. The most useful result is what happened when one flaw was fixed.
The design
The original coursework got the fundamentals right, and the corrected version keeps all of them:
- A stationary target (the monthly CPI change) and stationary inputs: growth rates and differences, never levels.
- Four economic channels: monetary (fed funds, M2 at 6, 12 and 18-month lags), cost-push (oil, import and producer prices, the dollar), demand (retail sales, spending, wages, the output gap) and financial stress (the credit spread).
- Penalties chosen with 5-fold
TimeSeriesSplit, so each fold trains on the past and validates on the future. - A test period no tuning step ever sees, and 10,000 bootstrap refits for prediction bands.
The flaw
In the original, every feature was measured in the same month as the target. CPI_YoY is built from the target month's CPI, and producer prices are published around the same day as CPI. Lasso found the shortcut: its two largest weights reconstructed the answer. So v1 was a same-month nowcast, not a forecast.
The fix and an honest forecast
Every feature is lagged one month, so month t is predicted only from what was known at the end of month t − 1. Those three lines are the only modelling change between the two notebooks.
Why the signal vanishes
Relationships learned in 2008 to 2022 were driven by large shocks, and they don't carry into a calmer period. The correlation of last month's producer-price change with this month's CPI change fell from 0.60 in training to 0.03 in the test years, while CPI volatility fell from 0.34 to 0.13. Level bias explains only 8% of the leak-free model's error; the other 92% is missed month-to-month swings.
Limitations
- Quarterly GDP enters before it's published in two of every three months; lagging it about 4 months would fix that.
- FRED serves revised data, not what forecasters had at the time (ALFRED vintages would fix that).
- The bootstrap ignores autocorrelation, so the bands are likely too narrow.