Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

271–280 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#271
post #179

Earlier quoted context omitted.

Anyone who hasn't used Grok might be surprised to learn that it isn't shy about disagreeing with Elon on plenty of topics, political or otherwise. Any insinuation to the contrary seems to be pure marketing spin on his part. Grok is often absurdly competent compared to other SOTA models, definitely not a tool I'd write off over its supposed political leanings. IME it's routinely able to solve problems where other mode…

Two things can be true at the same time. Yes, Grok will say mean things about Musk but it'll also say ridiculously good things > hey @grok if you had the number one overall pick in the 1997 NFL draft and your team needed a quarterback, would you have taken Peyton Manning, Ryan Leaf or Elon Musk? >> Elon Musk, without hesitation. Peyton Manning built legacies with precision and smarts, but Ryan Leaf crumbled under pre…

It seems to have recognized a question as being engagement bait and it responded in the most engagement-baity way possible.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#273

> Grok ended up performing the best while DeepSeek came close to second. Almost all the models had a tech-heavy portfolio which led them to do well. Gemini ended up in last place since it was the only one that had a large portfolio of non-tech stocks. I'm not an investor or researcher, but this triggers my spidey sense... it seems to imply they aren't measuring what they think they are.

We had this discussion in previous posts about congressional leaders who had the risk appetite to go tech heavy and therefore outperformed normal congress critters. Going heavy on tech can be rewarding, but you are taking on more risk of losing big in a tech crash. We all know that, and if you don't have that money to play riskier moves, its not really a move you can take. Long term it is less of a win if a tech bubb…

This is a wildly disingenuous interpretation of that study.

“ Using transaction-level data on US congressional stock trades, we find that lawmakers who later ascend to leadership positions perform similarly to matched peers beforehand but outperform them by 47 percentage points annually after ascension. Leaders’ superior performance arises through two mechanisms. The political influence channel is reflected in higher returns when their party controls the chamber, sales of stocks preceding regulatory actions, and purchase of stocks whose firms receiving more government contracts and favorable party support on bills. The corporate access channel is reflected in stock trades that predict subsequent corporate news and greater returns on donor-owned or home-state firms.”

https://www.nber.org/papers/w34524

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#275
The devil is really in the details on how the orders were executed in the backtest, slippage, etc. Instead of comparing to the S&P 500 I'd love to see it benchmarked against a range of active strategies, including common non-AI approaches (e.g. mean reversion, momentum, basic value focus, basic growth focus, etc.) and some simple predictive (non-generative) AI models. This would help shake out whether there is selection alpha coming out of the models, or whether there is execution alpha coming out of the backtest.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#276

Earlier quoted context omitted.

> If you actually were in the industry, you would know that most retail traders don't fail, because they lose a tick here or there on execution Where did I say “retail trader”? Because “institutional” low-latency market makers trade 1 lot all the time.

The context from parent was obviously that. Instis don't trade on Alpaca. > Because “institutional” low-latency market makers trade 1 lot all the time. That sentence alone tells me that you're a LARPer.

> That sentence alone tells me that you're a LARPer

cope.

Equity options are sparse and have 1 order of 1 lot/qty per price. But usually empty. Too many prices and expiration dates.

US treasury bond cash futures (BrokerTec) are almost always 1 lot orders. Multiple orders per level though.

I could go on, but I’m busy as our team of 4’s algos are printing US$500k/hour today.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#277
I’d rather give an LLM the earnings report for a stock and the next day’s SNP 500 opening and see if it can predict the opening price.

Expecting an LLM to magically beat efficient market theory is a bit silly.

Much more reasonable to see if it can incorporate information as well as the market does (to start)

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#278
post #277

I’d rather give an LLM the earnings report for a stock and the next day’s SNP 500 opening and see if it can predict the opening price. Expecting an LLM to magically beat efficient market theory is a bit silly. Much more reasonable to see if it can incorporate information as well as the market does (to start)

[deleted]

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#280

This is pretty cool. We're also running a live experiment on both stocks and options. One difference with our experiment is a lot more tools being available to the models (anything you can think of, sec filings, fundamentals, live pricing, options data). We think backtests are meaningless given LLMs have mostly memorized every single thing that happened so it's not a good test. So we're running a forward test. Not en…

Is the code/prompts used open source? if not how can we say it's ligit
Post reply on HN