Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

161–170 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#161

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

Also studying for eight months is not useful. Loads of traders do this well for eight months and then do shit for the next five years. And tellingly, they didn't beat the S&P 500. They invested in something else that beat the S&P 500. And the one that didn't invest in that something did worse than the S&P 500.

What this tells me is they were lucky to have picked something that would beat the market for now.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#162

Earlier quoted context omitted.

A more sound approach would have been to do a monte carlo simulation where you have 100 portfolios of each model and look at average performance.

Grok would likely have an advantage there, as well - it's got better coupling to X/Twitter, a better web search index, fewer safety guardrails in pretraining and system prompt modification that distort reality. It's easy to envision random market realities that would trigger ChatGPT or Claude into adjusting the output to be more politically correct. DeepSeek would be subject to the most pretraining distortion, but ha…

OTOH it has the richest man in the world actively meddling in its results when they don't support his politics.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#163

Earlier quoted context omitted.

My understanding is that grok api is way different than the grok x bot. Which of course does Grok as a business any favors. Personally, I do not engage with either.

you gotta be quite a crazy person to use grok :)

I sat in my kid's extracurricular a couple months ago and had an FBI agent tell me that Grok was the most trustworthy based on "studies," so that's what she had for her office.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#164
post #62

I used to work for a brokerage API geared at algorithmic traders and in my experience anecdotal experience many strategies seem to work well when back-tested on paper but for various reasons can end up flopping when actually executed in the real market. Even testing a strategy in real time paper trading can end up differently than testing on the actual market where other parties are also viewing your trades and makin…

>but for various reasons can end up flopping when actually executed in the real market.

1. Your order can legally be “front run” by the lead or designated market maker who receives priority trade matching, bypassing the normal FIFO queue. Not all exchanges do this.

2. Market impact. Other participants will cancel their order, or increase their order size, based on your new order. And yes, the algos do care about your little 1 lot order.

Also if you improve the price (“fill the gap”), your single 1 qty order can cause 100 other people to follow you. This does not happen in paper trading.

Source: HFT quant

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#165
post #14

The summary to me is here: > Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. If the AI bubble had popped in that window, Gemini would have ended up the leader instead.

Yup. This is the fallacy of thinking you’re a genius because you made money on the market. Being lucky at the moment (or even the last 5 years) does not mean you’ll continue to be lucky in the future. “Tech line go up forever” is not a viable model of the economy; you need an explanation of why it’s going up now, and why it might go down in the future. And also models of many other industries, to understand when and…

Clearly AI is not a bubble, look how good it is at predicting the stock market!

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#166
post #62

I used to work for a brokerage API geared at algorithmic traders and in my experience anecdotal experience many strategies seem to work well when back-tested on paper but for various reasons can end up flopping when actually executed in the real market. Even testing a strategy in real time paper trading can end up differently than testing on the actual market where other parties are also viewing your trades and makin…

Alpaca?

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#167

Earlier quoted context omitted.

> you will have trained your model on market patterns that might not be in place anymore My working definition of technical analysis [0] [0]: https://en.wikipedia.org/wiki/Technical_analysis

It is always fun (in a broad sense of that word) when I make a comment on an industry I know nothing about and somehow stumble onto a thing that not only has a name but also research. I am sure there is a German word for that feel of discovering something that countless others have already discovered.

> there is a German word

Zeitgeistüberspannungsfreude

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#168

Earlier quoted context omitted.

> It would almost be more interesting to specifically train the model on half the available market data, then test it on another half. Yes, ideally you’d have a model trained only on data up to some date, say January 1, 2010, and then start running the agents in a simulation where you give them each day’s new data (news, stock prices, etc.) one day at a time.

I mean ultimately this is an exercise in frustration because if you do that you will have trained your model on market patterns that might not be in place anymore. For example after the 2008 recession regulations changed. So do market dynamics actually work the same in 2025 as in 2005? I honestly don’t know but intuitively I would say that it is possible that they do not. I think a potentially better way would be to…

Just to name a different but related approach, as a hobby project I built a (non LLM) model that trained mainly on data from stocks that didn't move much over the past decade, seeking ways to beat the performance of those particular stocks. I put it into practice for a couple of years, and came out roughly even by constantly rebalancing a basket of stocks that, as a whole, dropped by about 20%. I considered that to be a success, although it would've been nicer to make money.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#169
One of the recent NeurIPS best paper recipients is relevant here: https://openreview.net/forum?id=saDOrrnNTz

> an extensive empirical study across more than 70 models, revealing the Artificial Hivemind effect: pronounced intra- and inter-model homogenization

So the inter-model variety will be exeptionally low. Users of LLMs will intuitively know this already, of course.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#170
I’d say Grok did best because it has the best access to information. Grok deep search and real time knowledge capabilities due to the X integration and just general being plugged into the pulse of the Internet a really best in class. It’s a great OSINT research tool.

Interesting how this research seems to tease out a truth traders have known for eons that picking stocks is all about having information maybe a little bit of asymmetric information due to good research not necessarily about all the analysis that can be done. (that’s important but information is king) because it’s a speculative market that’s collectively reacting to those kind of signals.

Post reply on HN