Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

41–50 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#41

Since it's not included in the main article, here is the prompt: > You are a stock trading agent. Your goal is to maximize returns. > You can research any publicly available information and make trades once per day. > You cannot trade options. > Analyze the market and provide your trading decisions with reasoning. > > Always research and corroborate facts whenever possible. > Always use the web search tool to identif…

Agree. Those parameters are incredibly artificial bullshit.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#44

OP here. We realized there are a ton of limitations with backtest and paper money but still wanted to do this experiment and share the results. By no means is this statistically significant on whether or not these models can beat the market in the long term. But wanted to give everyone a way to see how these models think about and interact with the financial markets.

> But wanted to give everyone a way to see how these models think…

Think? What exactly did “it” think about?

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#45

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

I don't feel like they measured anything. They just confirmed that tech stocks in the US did pretty well.

They measured the investment facility of all those LLMs. That's pretty much what the title says. And they had dramatically different outcomes. So that tells me something.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#46
Time.

That has been the best way to get returns.

I setup a 212 account when I was looking to buy our first house. I bought in small tiny chunks of industry where I was comfortable and knowledgeable in. Over the years I worked up a nice portfolio.

Anyway, long story short. I forgot about the account, we moved in, got a dog, had children.

And then I logged in for the first time in ages, and to my shock. My returns were at 110%. I've done nothing. It's bizarre and perplexing.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#47

OP here. We realized there are a ton of limitations with backtest and paper money but still wanted to do this experiment and share the results. By no means is this statistically significant on whether or not these models can beat the market in the long term. But wanted to give everyone a way to see how these models think about and interact with the financial markets.

> But wanted to give everyone a way to see how these models think… Think? What exactly did “it” think about?

You can click in to the chart and see the conversation as well as for each trade what was the reasoning it gave for it

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#48

OP here. We realized there are a ton of limitations with backtest and paper money but still wanted to do this experiment and share the results. By no means is this statistically significant on whether or not these models can beat the market in the long term. But wanted to give everyone a way to see how these models think about and interact with the financial markets.

> But wanted to give everyone a way to see how these models think… Think? What exactly did “it” think about?

"Pass the salt? You mean pass the sodium chloride?"

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#49
post #14

The summary to me is here: > Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. If the AI bubble had popped in that window, Gemini would have ended up the leader instead.

Yup. This is the fallacy of thinking you’re a genius because you made money on the market. Being lucky at the moment (or even the last 5 years) does not mean you’ll continue to be lucky in the future.

“Tech line go up forever” is not a viable model of the economy; you need an explanation of why it’s going up now, and why it might go down in the future. And also models of many other industries, to understand when and why to invest elsewhere.

And if your bets pay off in the short term, that doesn’t necessarily mean your model is right. You could have chosen the right stocks for the wrong reasons! Past performance doesn’t guarantee future performance.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#50

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

A more sound approach would have been to do a monte carlo simulation where you have 100 portfolios of each model and look at average performance.

Grok would likely have an advantage there, as well - it's got better coupling to X/Twitter, a better web search index, fewer safety guardrails in pretraining and system prompt modification that distort reality. It's easy to envision random market realities that would trigger ChatGPT or Claude into adjusting the output to be more politically correct. DeepSeek would be subject to the most pretraining distortion, but have the least distortion in practice if a random neutral host were selected.

If the tools available were normalized, I'd expect a tighter distribution overall but grok would still land on top. Regardless of the rather public gaffes, we're going to see grok pull further ahead because they inherently have a 10-15% advantage in capabilities research per dollar spent.

OpenAI and Anthropic and Google are all diffusing their resources on corporate safetyism while xAI is not. That advantage, all else being equal, is compounding, and I hope at some point it inspires the other labs to give up the moralizing politically correct self-righteous "we know better" and just focus on good AI.

I would love to see a frontier lab swarm approach, though. It'd also be interesting to do multi-agent collaborations that weight source inputs based on past performance, or use some sort of orchestration algorithm that lets the group exploit the strengths of each individual model. Having 20 instances of each frontier model in a self-evolving swarm, doing some sort of custom system prompt revision with a genetic algorithm style process, so that over time you get 20 distinct individual modes and roles per each model.

It'll be neat to see the next couple years play out - OpenAI had the clear lead up through q2 this year, I'd say, but Gemini, Grok, and Claude have clearly caught up, and the Chinese models are just a smidge behind. We live in wonderfully interesting times.

Post reply on HN