Earlier quoted context omitted.
A more sound approach would have been to do a monte carlo simulation where you have 100 portfolios of each model and look at average performance.
Grok would likely have an advantage there, as well - it's got better coupling to X/Twitter, a better web search index, fewer safety guardrails in pretraining and system prompt modification that distort reality. It's easy to envision random market realities that would trigger ChatGPT or Claude into adjusting the output to be more politically correct. DeepSeek would be subject to the most pretraining distortion, but ha…
Really? Isn't Grok's whole schtick that it's Elon's personal altipedia?