Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

111–120 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#111

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

We had this discussion in previous posts about congressional leaders who had the risk appetite to go tech heavy and therefore outperformed normal congress critters.

Going heavy on tech can be rewarding, but you are taking on more risk of losing big in a tech crash. We all know that, and if you don't have that money to play riskier moves, its not really a move you can take.

Long term it is less of a win if a tech bubble builds and pops before you can exit (and you can't out it out to re-inflate).

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#112
post #55

Earlier quoted context omitted.

They measured the investment facility of all those LLMs. That's pretty much what the title says. And they had dramatically different outcomes. So that tells me something.

I mean, what it kinda tells me is that people talk about tech stocks the most, so that's what was most prevalent in the training data, so that's what most of the LLMs said to invest in. That's the kind of strategy that works until it really doesn't.

Cue 2020 or so. I do have investments in tech stocks but I have a lot more conservative investments too.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#113
post #10

There's also this thing going on right now: https://nof1.ai/leaderboard Results are... underwhelming. All the AIs are focused on daytrading Mag7 stocks; almost all have lost money with gusto.

I also saw the hype on X yesterday and had already checked the https://nof1.ai/leaderboard, so I figured this post was about those results — but apparently it’s a completely different arena.

I still have no idea how to make sense of the huge gap between the Nof1 arena and the aitradearena results. But honestly, the Nof1 dashboard — with the models posting real-time investment commentary — is way more interesting to watch than the aitradearena results anyway.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#115
I setup real life accounts with etrade and fidelity using the etrade auto portfolio, fidelity i have an advisor for retirement, and then i did a basket portfolio as well but used ms365 with grok 5 and various articles and strategies to pick a set of 5 etfs that would perform similarly to the exposure of my other two.

This year So far all are beating the s&p % wise (only by In other words though im not surprised at all by the results. Ai isnt something to day trade with still but it is helpful in doing research for your desired risk exposure long term imo.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#116

Just one run per model? That isn't backtesting. I mean technically it is, but "testing" implies producing meaningful measures. Also just one time interval? Something as trivial as "buy AI" could do well in one interval, and given models are going to be pumped about AI, ... 100 independent runs on each model over 10 very different market behavior time intervals would producing meaningful results. Like actually credibl…

To their credit, they say in the article that the results aren't statistically significant. It would be better if that disclaimer was more prominently displayed though.

The tone of the article is focused on the results when it should be "we know the results are garbage noise, but here is an interesting idea".

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#117

Earlier quoted context omitted.

I mean ultimately this is an exercise in frustration because if you do that you will have trained your model on market patterns that might not be in place anymore. For example after the 2008 recession regulations changed. So do market dynamics actually work the same in 2025 as in 2005? I honestly don’t know but intuitively I would say that it is possible that they do not. I think a potentially better way would be to…

> you will have trained your model on market patterns that might not be in place anymore My working definition of technical analysis [0] [0]: https://en.wikipedia.org/wiki/Technical_analysis

It is always fun (in a broad sense of that word) when I make a comment on an industry I know nothing about and somehow stumble onto a thing that not only has a name but also research. I am sure there is a German word for that feel of discovering something that countless others have already discovered.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#118

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

I mean, run the experiment during a different trend in the market and the results would probably be wildly different. This feels like chartists [1] but lazier.

[1] https://www.investopedia.com/terms/c/chartist.asp

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#119
Just picking tech stocks and winning isn't interesting unless we know the thesis behind picking the tech sticks.

Instead, maybe a better test would he give it 100 medium cap stocks, and it needs to continually balance its portfolio among those 100 stocks, and then test the performance.

Post reply on HN