Im curious, how do data scientists use these massive datasets, especially the old stuff. Is it more of a compliance and need/should-save type thing or is the data actually useful? Im baffled by these numbers having never used a large BI tool, and am genuinely curious how the data is actually used operationally. As a layman, I imagine lots of it loses relevancy very quickly, e.g Amazon sales data from 5 years ago is m…
I work in finance and it's great having big historical datasets, even if the figures are far lower in previous years it's good to see system 'shocks' and these can be used at a different magnitude/scaled for future forecasting
Related, I rather enjoyed reading this other thread from June: "Ask HN: Is KDB a sane choice for a datalake in 2024?" https://news.ycombinator.com/item?id=40625800