Im pretty sure its just 4.0 but it re-prompts itself a few times before answering. It costs a lot more
Seems like Reflection 70b was an attempt to implement the same concept on top of Llama 3 70b
Pushing an hypothetical (and likely false, but not impossible) conspiracy theory much further:
in theory, they had access in their backend logs to the prompts that Reflection 70b were doing while calling GPT-4o (as it apparently was actually calling both Anthropic and OpenAI API instead of LLaMA), and had an opportunity to get "inspired".