Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

281–290 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#281

Earlier quoted context omitted.

>but for various reasons can end up flopping when actually executed in the real market. 1. Your order can legally be “front run” by the lead or designated market maker who receives priority trade matching, bypassing the normal FIFO queue. Not all exchanges do this. 2. Market impact. Other participants will cancel their order, or increase their order size, based on your new order. And yes, the algos do care about your…

Dear HFT Quant, > And yes, the algos do care about your little 1 lot order. I'm just your usual "corrupted nerd" geek with some mathematics and computer security background interests - 2 questions if I may 1. what's like the most interesting paper you have read recently or unrelated thing you are interested in at the moment? 2. " And yes, the algos do care about your little 1 lot order." How would one see this effect…

Even a 1 lot order could be the deciding factor for some algorithm that's calculating averages or other statistics. Especially for options books.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#282
post #163

Earlier quoted context omitted.

you gotta be quite a crazy person to use grok :)

I sat in my kid's extracurricular a couple months ago and had an FBI agent tell me that Grok was the most trustworthy based on "studies," so that's what she had for her office.

Grok has Elon as better athelete than LeBron so I would agree with FBI Agent. can’t get that kind of insight anywhere else :)

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#284

Earlier quoted context omitted.

>but for various reasons can end up flopping when actually executed in the real market. 1. Your order can legally be “front run” by the lead or designated market maker who receives priority trade matching, bypassing the normal FIFO queue. Not all exchanges do this. 2. Market impact. Other participants will cancel their order, or increase their order size, based on your new order. And yes, the algos do care about your…

>Your order can legally be “front run” by the lead or designated market maker who receives priority trade matching, bypassing the normal FIFO queue. Not all exchanges do this. Unless you're thinking of some obscure exchange in a tiny market, this is just untrue in the U.S., Europe, Canada, and APAC. There are no exchanges where market makers get any kind of priority to bypass the FIFO queue.

> There are no exchanges where market makers get any kind of priority to bypass the FIFO queue.

Nope, several large, active, and liquid markets in the US.

Legally it’s not named “bypass the FIFO queue”. That would be dumb.

In practice, it goes by politically correct names such as “designated market maker fill” or “institutional order prioritization” or “leveling round”.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#285
This is the complete wrong way to do this. I say this as someone who does work in this area of leveraging LLMs to a limited degree in trading.

LLMs are naive, easily convinced, and myopic. They're also non-deterministic. We have no way of knowing if you ran this little experiment 10 times whether they'd all pick something else. This is a scattershot + luck.

The RIGHT way to do this is to first solve the underlying problem deterministically. That is, you first write your trading algorithm that's been thoroughly tested. THEN you can surface metadata to LLMs and say things along the lines of "given this data + data you pull from the web", make your trade decision for this time period and provide justification.

Honestly, adding LLMs directly to any trading pipeline just adds non-useful non-deterministic behavior.

The main value is speed of wiring up something like sentiment analysis as a value add or algorithmic supplement. Even this should be done using proper ML but I see the most value in using LLMs to shortcut ML things that would require time/money/compute. Trading value now for value later (the ML algorithm would ultimately run cheaper long-run but take longer to get into prod).

This experiment, like most "I used AI to trade" blogs are completely naive in their approach. They're taking the lowest possible hanging fruit. Worst still when those results are the rising tide lifting all boats.

Edit (was a bit harsh) This experiment is an example of the kind of embarrassingly obvious things people try with LLMs without understanding the domain and writing it up. To an outsider it can sound exciting. To an insider it's like seeing a new story "LLMs are designing new CPUs!". No they're not. A more useful bit of research would be to control for the various variables (sector exposure etc) and then run it 10_000 times and report back on how LLM A skews towards always buying tech and LLM B skews towards always recommending safe stocks.

Alternatively, if they showed the LLM taking a step back and saying "ah, let me design this quant algo to select the best stocks" -- and then succeeding -- I'd be impressed. I'd also know that it was learned from every quant that had AI double check their calculations/models/python.. but that's a different point.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#286

Earlier quoted context omitted.

They measured the investment facility of all those LLMs. That's pretty much what the title says. And they had dramatically different outcomes. So that tells me something.

It shows nothing. This is a bullshit stunt that should be obvious to anyone who has placed a few trades.

Unless you think of it as an AI exercise, not a stock trading exercise. Which point evaded most people.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#287

Earlier quoted context omitted.

They measured the investment facility of all those LLMs. That's pretty much what the title says. And they had dramatically different outcomes. So that tells me something.

They "proved" that US tech stocks did better than portfolios with less US tech stocks over a recent, very short time range. 1. You didn't know that? 2. Whata re you going to do with this "new information"?

As a stock-trading exercise? Nothing, as you note. As an AI investigation it says plenty. Which is the point I was making (and got missed by all those stock-trading self-appointed experts who fastened onto that)

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#288
post #62

I used to work for a brokerage API geared at algorithmic traders and in my experience anecdotal experience many strategies seem to work well when back-tested on paper but for various reasons can end up flopping when actually executed in the real market. Even testing a strategy in real time paper trading can end up differently than testing on the actual market where other parties are also viewing your trades and makin…

Backtracking is useless because if you try out a million strategies, by chance you will find one that works for past data.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#289

Earlier quoted context omitted.

When does the tech sector become the computer sector? Agriculture would have been considered tech 200 years ago.

full throttle until AGI is achieved, then we will see

Maybe one day we will discover that a method exists for computing/displaying/exchanging arbitrary things through none other means than our own flesh and brains.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#290

Earlier quoted context omitted.

Grok would likely have an advantage there, as well - it's got better coupling to X/Twitter, a better web search index, fewer safety guardrails in pretraining and system prompt modification that distort reality. It's easy to envision random market realities that would trigger ChatGPT or Claude into adjusting the output to be more politically correct. DeepSeek would be subject to the most pretraining distortion, but ha…

> fewer safety guardrails in pretraining and system prompt modification that distort reality. Really? Isn't Grok's whole schtick that it's Elon's personal altipedia?

It's excellent, and it doesn't get into the weird ideological ruts and refusals other bots do.

Grok's search and chat is better than the other platforms, but not $300/month better, ChatGPT seems to be the best no rate limits pro class bot. If Grok 5 is a similar leap in capabilities as 3 to 4, then I might pay the extra $100 a month. The "right wing Elon sycophant" thing is a meme based on hiccups with the public facing twitter bot. The app, api, and web bot are just generally very good, and do a much better job at neutrality and counterfactuals and not refusing over weird moralistic nonsense.

Post reply on HN