Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

11–20 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#12
> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks.

I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#13
I wonder if this could be explained as the result of LLMs being trained to have pro-tech/ai opinions while we see massive run ups in tech stock valuations?

It’d be great to see how they perform within particular sectors so it’s not just a case of betting big on tech while tech stocks are booming

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#14
The summary to me is here:

> Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks.

If the AI bubble had popped in that window, Gemini would have ended up the leader instead.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#16
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

That's only if they're trained on data more recent than 8 months ago

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#18

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

I don't feel like they measured anything. They just confirmed that tech stocks in the US did pretty well.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#19

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

A more sound approach would have been to do a monte carlo simulation where you have 100 portfolios of each model and look at average performance.
Post reply on HN