Live data from Hacker News

We gave 5 LLMs $100K to trade stocks for 8 months

aitradearena.com

1–10 of 319 posts

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#3
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

Not sure how sound the analysis is but they did apparently actually think of that.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#4
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

> We were cautious to only run after each model’s training cutoff dates for the LLM models. That way we could be sure models couldn’t have memorized market outcomes.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#5
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

> We time segmented the APIs to make sure that the simulation isn’t leaking the future into the model’s context.

I wish they could explain what this actually means.

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#6
post #3
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

Not sure how sound the analysis is but they did apparently actually think of that.

[deleted]

Re: We gave 5 LLMs $100K to trade stocks for 8 months

#9
post #4
post #2

> Testing GPT-5, Claude, Gemini, Grok, and DeepSeek with $100K each over 8 months of backtested trading So the results are meaningless - these LLMs have the advantage of foresight over historical data.

> We were cautious to only run after each model’s training cutoff dates for the LLM models. That way we could be sure models couldn’t have memorized market outcomes.

I know very little about how the environment where they run these models look, but surely they have access to different tools like vector embeddings with more current data on various topics?
Post reply on HN