I don’t really understand why Google and other companies making similar models are able to train on existing modelled or reanalysis data sets and then claim further accuracy than the originals. Sure, stacks of convolutions with multimodal attention blocks should be able to tease apart all the of idiosyncratic correlations that the original models may not have seen. But it’s unclear to me that better models is the dir…
> The model was trained on four decades of temperature, wind speed and air pressure data from 1979 to 2018 and can produce a 15-day forecast in just eight minutes—compared to the hours it currently takes. Basically, they trained the model on old observations, not old predictions. IE, imagine someone created a giant spreadsheet of every temperature observation, windspeed observation, precipitation observation, ect, an…
> GenCast is trained on 40 years of best-estimate analysis from 1979 to 2018, taken from the publicly available ERA5 (fifth generation ECMWF reanalysis) reanalysis dataset
https://www.nature.com/articles/s41586-024-08252-9
They say they trained on era5 data. That is modeled data not only direct observations. They do it this way to have a complete global data set that is consistent in time and space.
https://cds.climate.copernicus.eu/datasets/reanalysis-era5-s...