3 Statsmodels Tricks for Time Series Analysis & Forecasting

A fitted statsmodels model computes a more than just the array of numbers most code pulls out of it. The point forecast is the smallest task it can perform. Every trick runs from a case of asking the results object for something it has already worked out, rather than having to rebuild that thing by hand. One dataset, one model, three methods people routinely reimplement. All three are run against the same monthly series and the same fitted model, so the only thing that changes between them is which method gets called on the object fit handed back.
Everything below was checked against statsmodels 0.15.0.
Start by installing statsmodels:
pip install statsmodels
Trick 1: Asking for the Interval, Not Just the Number
res.forecast(12) gives you twelve numbers. res.get_forecast(12) gives you a PredictionResults object instead, which happens to also carry the uncertainty the model already estimated. There’s predicted_mean for the points, conf_int for the bounds, and summary_frame() for both at once. The intervals aren’t even extra work; they are a result of the same computation, and the shorter method simply throws them away:
import statsmodels.api as sm
from statsmodels.tsa.arima.model import ARIMA
co2 = sm.datasets.co2.load_pandas().data["co2"]
co2 = co2.resample("MS").mean().ffill()
train, recent = co2[:-12], co2[-12:]
res = ARIMA(train, order=(1, 1, 1), seasonal_order=(1, 1, 1, 12)).fit()
print(res.get_forecast(12).summary_frame().head())
Output:
co2 mean mean_se mean_ci_lower mean_ci_upper
2001-01-01 370.523929 0.322722 369.891406 371.156452
2001-02-01 371.253673 0.388214 370.492787 372.014559
2001-03-01 372.200726 0.429518 371.358887 373.042566
2001-04-01 373.468351 0.463501 372.559905 374.376797
2001-05-01 373.856957 0.494157 372.888427 374.825487
Publishing a forecast without its interval is a choice. And if what you want is fitted values over the history rather than the path ahead, get_prediction is this same idea applied to a range that can include in-sample periods.
Trick 2: Adding New Data Without Refitting
Twelve months of observations arrive. The reflex is to concatenate them onto the training data and call .fit() again, which re-estimates every parameter from scratch. append does something cheaper: it recreates the results object over the combined data and, with refit=False, it keeps the parameters you already estimated:
updated = res.append(recent, refit=False)
print(updated.get_forecast(6).summary_frame().head())
Output:
co2 mean mean_se mean_ci_lower mean_ci_upper
2002-01-01 371.969954 0.322722 371.337432 372.602477
2002-02-01 372.750021 0.388214 371.989135 373.510907
2002-03-01 373.654908 0.429518 372.813068 374.496748
2002-04-01 374.834276 0.463501 373.925830 375.742722
2002-05-01 375.328719 0.494157 374.360189 376.297248
The default is refit=False, which reuses the estimates you already have. Pass refit=True when enough new data has accumulated that you want them computed again.
There are three of these methods:
appendre-runs the filter over the original data as well as the newextendfilters only the new observations, which is faster when the history is longapplyis for a different dataset rather than a continuation of this one
Trick 3: Letting STL Do the Seasonality
The manual version of this procedure is three steps:
- decompose the series
- forecast the seasonally adjusted part
- then add the seasonal component back onto the forecast
It’s the third step where the sign errors and the index misalignments happen. STLForecast is that whole loop as one object. The docs describe it as forecasting “by first subtracting the seasonality estimated using STL, then forecasting the deseasonalized data using a time-series model, for example, ARIMA”:
from statsmodels.tsa.forecasting.stl import STLForecast
stlf = STLForecast(train, ARIMA, model_kwargs={"order": (1, 1, 1), "trend": "t"})
print(stlf.fit().forecast(12).head())
Output:
2001-01-01 370.529117
2001-02-01 370.963627
2001-03-01 371.921080
2001-04-01 373.137720
2001-05-01 373.144192
Freq: MS, dtype: float64
Note what gets passed: the ARIMA class itself, not a fitted instance, with its arguments handed over separately in model_kwargs. That’s the one genuinely surprising thing in this API, and passing ARIMA(...) instead is the first mistake most people make with it.
Wrapping Up
Every trick here is a method that already exists on an object you have already built. The hand-rolled alternative is longer, slower, and wrong more often, so deviating from the built-ins in this instance is genuinely not worth it. It usually gets written because nobody looked at what came back from .fit(). Read the results object; then stop rewriting it.
Matthew Mayo (@mattmayo13) holds a master’s degree in computer science and a graduate diploma in data mining. As managing editor of KDnuggets & Statology, and contributing editor at Machine Learning Mastery, Matthew aims to make complex data science concepts accessible. His professional interests include natural language processing, language models, machine learning algorithms, and exploring emerging AI. He is driven by a mission to democratize knowledge in the data science community. Matthew has been coding since he was 6 years old.



