We gave 5 LLMs $100K to trade stocks for 8 months
121–130 of 319 posts
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#122Earlier quoted context omitted.
> you will have trained your model on market patterns that might not be in place anymore My working definition of technical analysis [0] [0]: https://en.wikipedia.org/wiki/Technical_analysis
It is always fun (in a broad sense of that word) when I make a comment on an industry I know nothing about and somehow stumble onto a thing that not only has a name but also research. I am sure there is a German word for that feel of discovering something that countless others have already discovered.
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#123Earlier quoted context omitted.
A more sound approach would have been to do a monte carlo simulation where you have 100 portfolios of each model and look at average performance.
Grok would likely have an advantage there, as well - it's got better coupling to X/Twitter, a better web search index, fewer safety guardrails in pretraining and system prompt modification that distort reality. It's easy to envision random market realities that would trigger ChatGPT or Claude into adjusting the output to be more politically correct. DeepSeek would be subject to the most pretraining distortion, but ha…
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#124Re: We gave 5 LLMs $100K to trade stocks for 8 months
#125> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#126Earlier quoted context omitted.
My understanding is that grok api is way different than the grok x bot. Which of course does Grok as a business any favors. Personally, I do not engage with either.
you gotta be quite a crazy person to use grok :)
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#127Just one run per model? That isn't backtesting. I mean technically it is, but "testing" implies producing meaningful measures. Also just one time interval? Something as trivial as "buy AI" could do well in one interval, and given models are going to be pumped about AI, ... 100 independent runs on each model over 10 very different market behavior time intervals would producing meaningful results. Like actually credibl…
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#128It turns out DeepSeek only made BUY trades (not a single SELL in the history in their live example) -- so basically, buy & hold strategy wins, again.
Re: We gave 5 LLMs $100K to trade stocks for 8 months
#129Re: We gave 5 LLMs $100K to trade stocks for 8 months
#130> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.
We had this discussion in previous posts about congressional leaders who had the risk appetite to go tech heavy and therefore outperformed normal congress critters. Going heavy on tech can be rewarding, but you are taking on more risk of losing big in a tech crash. We all know that, and if you don't have that money to play riskier moves, its not really a move you can take. Long term it is less of a win if a tech bubb…