> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.
I mean, run the experiment during a different trend in the market and the results would probably be wildly different. This feels like chartists [1] but lazier. [1] https://www.investopedia.com/terms/c/chartist.asp
We gave 5 LLMs $100K to trade stocks for 8 months
221–230 of 319 posts
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#222Re: We gave 5 LLMs $100K to trade stocks for 8 months
#223Earlier quoted context omitted.
this study should be replicated during a bear market
Buy and hold performs well over long time scales by simply not adjusting based upon sentiment.
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#224Earlier quoted context omitted.
Yeah I mean if you generally believe the tech sector is going to do well because it has been doing well you will beat the overall market. The problem is that you don’t know if and when there might be a correction. But since there is this one segment of the overall market that has this steady upwards trend and it hasn’t had a large crash, then yeah any pattern seeking system will identify “hey this line keeps going up…
> a hedge fund can beat the market for 2-4 years but at 10 years and up their chances of beating the market go to very close In that case the winning strategy would be to switch hedge funds every 3 years.
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#225> We also built a way to simulate what an agent would have seen at any point in the past. Each model gets access to market data, news APIs, company financials—but all time filtered: agents see only what would have been available on that specific day during the test period. That's not going to work, these agents especially the larger ones, will have news about the companies embedded in their weights.
> We were cautious to only run after each model’s training cutoff dates for the LLM models. That way we could be sure models couldn’t have memorized market outcomes.
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#226Earlier quoted context omitted.
OTOH it has the richest man in the world actively meddling in its results when they don't support his politics.
Anyone who hasn't used Grok might be surprised to learn that it isn't shy about disagreeing with Elon on plenty of topics, political or otherwise. Any insinuation to the contrary seems to be pure marketing spin on his part. Grok is often absurdly competent compared to other SOTA models, definitely not a tool I'd write off over its supposed political leanings. IME it's routinely able to solve problems where other mode…
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#227Earlier quoted context omitted.
you gotta be quite a crazy person to use grok :)
I sat in my kid's extracurricular a couple months ago and had an FBI agent tell me that Grok was the most trustworthy based on "studies," so that's what she had for her office.
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#228Earlier quoted context omitted.
I don't feel like they measured anything. They just confirmed that tech stocks in the US did pretty well.
They measured the investment facility of all those LLMs. That's pretty much what the title says. And they had dramatically different outcomes. So that tells me something.