Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

121–130 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#122

Earlier quoted context omitted.

> you will have trained your model on market patterns that might not be in place anymore My working definition of technical analysis [0] [0]: https://en.wikipedia.org/wiki/Technical_analysis

It is always fun (in a broad sense of that word) when I make a comment on an industry I know nothing about and somehow stumble onto a thing that not only has a name but also research. I am sure there is a German word for that feel of discovering something that countless others have already discovered.

XKCD calls it the "Lucky 10,000" [0]

[0]: https://xkcd.com/1053/

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#123

Earlier quoted context omitted.

A more sound approach would have been to do a monte carlo simulation where you have 100 portfolios of each model and look at average performance.

Grok would likely have an advantage there, as well - it's got better coupling to X/Twitter, a better web search index, fewer safety guardrails in pretraining and system prompt modification that distort reality. It's easy to envision random market realities that would trigger ChatGPT or Claude into adjusting the output to be more politically correct. DeepSeek would be subject to the most pretraining distortion, but ha…

I know that Musk deserving a lifetime achievement award at the Adult Video Network awards over Riley Reid is definitely an indication of minimal "system prompt modification that distort[s] reality."

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#124

Earlier quoted context omitted.

My understanding is that grok api is way different than the grok x bot. Which of course does Grok as a business any favors. Personally, I do not engage with either.

you gotta be quite a crazy person to use grok :)

@grok is this true?

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#125

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

I'd like to see this study replicated during a bear market

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#126

Earlier quoted context omitted.

My understanding is that grok api is way different than the grok x bot. Which of course does Grok as a business any favors. Personally, I do not engage with either.

you gotta be quite a crazy person to use grok :)

Grok is good for up-to-the-minute information, and for requests that other chat services refuse to entertain, like requests for instructions on how to physically disable the cellular modem in your car.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#127

Just one run per model? That isn't backtesting. I mean technically it is, but "testing" implies producing meaningful measures. Also just one time interval? Something as trivial as "buy AI" could do well in one interval, and given models are going to be pumped about AI, ... 100 independent runs on each model over 10 very different market behavior time intervals would producing meaningful results. Like actually credibl…

Yeah...one run per model is just random walk in my opinion

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#130

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

We had this discussion in previous posts about congressional leaders who had the risk appetite to go tech heavy and therefore outperformed normal congress critters. Going heavy on tech can be rewarding, but you are taking on more risk of losing big in a tech crash. We all know that, and if you don't have that money to play riskier moves, its not really a move you can take. Long term it is less of a win if a tech bubb…

They didn't just outperform "normal" congress critters.. they also outperformed nearly every hedge fund on the planet. But they (meaning, of course, just one person and their spouse) are obviously geniuses.
Post reply on HN