Earlier quoted context omitted.
You can find the entire Llama 3.0 pretraining set here: https://huggingface.co/datasets/HuggingFaceFW/fineweb 15T tokens, 45 terrabytes. Seems fairly open source to me.
Where has Facebook linked that? I can't find anywhere that they actually published that.
Regardless, it fits the compute used and the claim that they trained from public web data, and was suspiciously published by HF staff shortly after L3 released. It's about as official as the Mistral 7B v0.2 base model. I.e. mostly, but not entirely, probably for some weird legal reasons.