Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

21–30 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#22

>We gave each of five LLMs $100K in paper money Stopped reading after “paper money” Source: quant trader. paper trading does not incorporate market impact

If your initial portfolio is 100k you are not going to have meaningful "market impact" with your trades assuming you actually make them vs. paper trading.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#23
I think these tests are always difficult to gauge how meaningful they actually are. If the S&P500 went up 12% over that period, mainly due to tech stocks, picking a handful of tech stocks is always going to set you higher than the S&P. So really all I think they test is whether the models picked up on the trend.

I more surprised that Gemini managed to lose 10%. I wish they actually mentioned what the models invested in and why.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#24
post #5
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

> We time segmented the APIs to make sure that the simulation isn’t leaking the future into the model’s context. I wish they could explain what this actually means.

Overall, it does sound weird. On the one hand, assuming I properly I understand what they are saying is that they removed model's ability to cheat based on their specific training. And I do get that nuance ablation is a thing, but this is not what they are discussing there. They are only removing one avenue of the model to 'cheat'. For all we know, some that data may have been part of its training set already...

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#25

>We gave each of five LLMs $100K in paper money Stopped reading after “paper money” Source: quant trader. paper trading does not incorporate market impact

I mean if you’re going to write algos that trade the first thing you should do is check whether they were successful on historical data. This is an interesting data point.

Market impact shouldn’t be considered when you’re talking about trading S&P stocks with $100k.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#26
post #5
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

> We time segmented the APIs to make sure that the simulation isn’t leaking the future into the model’s context. I wish they could explain what this actually means.

It's a very silly way of saying that the data the LLMs had access to was presented in chronological order, so that for instance, when they were trading on stocks at the start of the 8 month window, the LLMs could not just query their APIs to see the data from the end of the 8 month window.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#28

>We gave each of five LLMs $100K in paper money Stopped reading after “paper money” Source: quant trader. paper trading does not incorporate market impact

Lack of market response is a valid point, but $100k is pretty unlikely to have much impact especially if spread out over multiple trades.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#29

Could they give some random people (i volunteer) 100k for 8 months? ...as a control

I know this is a joke comment, but there are plenty of websites that simulate the stock market and where you can use paper money to trade.

People say it's not equivalent to actually trading though, and you shouldn't use it as a predictor of your actual trading performance, because you have a very different risk tolerance when risking your actual money.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#30
post #11

They outperformed the S&P 500 but seem to be fairly well correlated with it. Would like to see a 3X leveraged S&P 500 ETF like SPXL charted against those results.

...over the course of 8.5 months, which is way too short for a meaningful result. If their strategy could outperform the S&P 500's 10-year return, they wouldn't be blogging about it.
Post reply on HN