Live data from Hacker News

The unreasonable difficulty of time series forecasting

suzyahyah.github.io

11–20 of 69 posts

Re: The unreasonable difficulty of time series forecasting

#12
This article focuses mostly on point forecasting. There are many other things that are interesting to forecast.

Consider for example the use case of forecasting the average speed on a road segment with a maximum speed of 70mph. Forecasting whether that will be 69.8 or 70.3 is not very relevant. What is relevant is forecasting when the speed drops below a traffic jam threshold. But the exact timing of that might be impossible to forecast due to the inherently chaotic behavior of traffic. Forecasting the probability of a traffic jam occurring may be more interesting to practitioners.

Re: The unreasonable difficulty of time series forecasting

#15
> Given the lack of forecasting signal, the obvious next step then is to seek out external (exogenous) features in the real world that can help prediction models.

this seems right to me. maybe another interesting approach would be a fusion llm+ts model that does multiple-input-single-output with input metadata and causality narrative. so it "thinks" about what data it has and how predictive it may be of the target variable and when something "interesting" occurs it uses the big priors to synthesize a good guess at what it would look like.

so you'd have something like the time series data plus textual narratives of the causality stories as the training data.

Re: The unreasonable difficulty of time series forecasting

#16
A very interesting article, suzyahyah. I especially appreciated your definition of stationarity, a concept with which I struggled in my own time series class. If I understand correctly, it sounds like the basic premise is that a fundamentally statistical methodology (LLMs) can't realistically predict a non-stationary data generation, which makes sense.

Separately, I've wondered for some time if there might be some reliable way to predict non-stationary data. While I don't have the answer, it occurs to me that it will possibly be a non-statistical method due to the fundamental incompatibilities. However, it also occurs to me that, given enough information, every data-generating process actually could be predicted. For instance, in the stock example, if you could model every single input into the system of a single company's stock, including every variable affecting every human that might conduct a transaction of it (daunting and unrealistic as that might be, but this is a thought experiment), then I believe the problem of prediction stops being non-stationary and in fact becomes completely deterministic, if complex. In such a scenario, wouldn't you be able to accurately make your prediction? I believe that perhaps chaos theory could present us with some solutions here where pure statistics (or, rather, simple statistics) cannot.

Just my 2 cents..

Re: The unreasonable difficulty of time series forecasting

#17
> Given the lack of forecasting signal, the obvious next step then is to seek out external (exogenous) features in the real world that can help prediction models. After all many real world time series are event-driven (FX, bitcoin) and hence they are exposed to shocks and drifts which can also be measured or accounted for in the data generating process.

This is the approach I'm using at my job, which is incident detection with customer metrics. We're tagging our time series data with common features -- such as country, customer type, etc -- with the idea that we can do a graph-like search to find exogenous variables. We can also use this to identify time series that have a similar "data generating process" and are simply different "realizations" of each other.

We don't need great time series forecasts, just something that detects large deviations quickly. We can then add in an existing dataset of _known_ incidents, indexed by the same common features, as a training/validation set.

Re: The unreasonable difficulty of time series forecasting

#18
post #4

I see that a lot of these are markets. Yes, it’s hard to predict markets. Because anybody who can successfully predict markets, does so, makes money, and changes the market so their predictions lose their edge. Time series forecasts are a lot easier if you are forecasting, say, disk use in your servers or whatnot. (By “easy” I mean you can do a simple prediction and get useful insights.)

It's true that markets are more _adversarial_. But there's still a lot of trouble with distribution shifts even in server metrics. As an example, our SRE team got paged a few times in the past month for traffic drops due to the World Cup. This stresses the nowcasting alert in several dimensions:

- there's no seasonal pattern to the matches, they happen sorta randomly.

- they drive increased query traffic in the hour or so before the game

- then during the game usage drops, sometimes to below "normal" depending on time of day and who's playing

So... now the accuracy of your forecasting tool depends on correctly predicting when world cup matches happen, and also who wins them!

edit: and this is just one recent example. others involve severe weather, national gameshows, earthquakes, and when you celebrate christmas.

Re: The unreasonable difficulty of time series forecasting

#19
post #4

I see that a lot of these are markets. Yes, it’s hard to predict markets. Because anybody who can successfully predict markets, does so, makes money, and changes the market so their predictions lose their edge. Time series forecasts are a lot easier if you are forecasting, say, disk use in your servers or whatnot. (By “easy” I mean you can do a simple prediction and get useful insights.)

Between the difficulty and value propositions for forecasting timeseries "disk usage" versus "financial markets" there are some rather relevant time series such as "company sales" or "demand for company product" that come up again, and again, and again but are neither "easy" nor "predicting-the-stock-market-hard".
Post reply on HN