Live data from Hacker News

A robot is sprinting towards you. Do you want it running on Claude or Grok?

openrouter.ai

121–130 of 234 posts

Re: A robot is sprinting towards you. Do you want it running on Claude or Grok?

#121

> I didn’t add any frontier-tier models like Opus 4.7, GPT-5.5, or Gemini Ultra. At their prices, 30 games would have cost around $3,000 instead of $482. I have a lot of thoughts unrelated to the game experiment but more about how these opus/ultra size models can possibly be a financially viable product at scale when it costs $3000 to play 30 simple games. It just seems much much higher than what it would cost to get…

> It just seems much much higher than what it would cost to get a human to play 30 rounds You mean almost like it was super short sighted to do a ton of layoffs when the AI tech is going to cost almost as much, if not more, than the humans it replaced? Yeah, you don't need Opus level for everything, and sonnet has gotten fairly decent I'm using it more and more, but still for most tasks I'm working with, Opus is the…

I experience the same with OpenAI, on the $100/month plan. GPT-5.4 is something I still have to challenge: it can bullshit me with bad implementation and add a lot of cruft that costs more time later. GPT-5.5-xhigh is something I have almost complete faith and trust in, it's just smooth. And yet I know the actual token cost of that fully utilized is exorbitant, like as much as an entire salary for a senior developer.

So maybe our CEOs are responding with a lot of foresight and inside information and know that that level of quality is going to be cheap really soon. But barring that, they're going to experience either sticker shock or a slowdown.

I think the real endgame is probably more accurate "models of models" (model routers) that know exactly how to split prompts between expensive frontier and cheap/free local models.

Re: A robot is sprinting towards you. Do you want it running on Claude or Grok?

#127

> I dropped eleven LLMs into a 2D battle royale and made them play 30 games. One won 43% of the matches. Three never won a single game. The cheapest model in the lineup beat the most expensive one by 27x on cost per win. Please learn how to write with AI without giving away that it was written by AI.

What about that makes you think it was written by AI?

The style is very obvious.

Some snippets that display classic patterns:

“ Both of those things are true. That’s the part most benchmarks can’t see,”

“And it’s changing how I” (classic pattern found in a lot of LinkedIn AIslop)

“ I want to be careful here.”

“ The stats are the stats. The moments are the part I kept showing people. ”

Re: A robot is sprinting towards you. Do you want it running on Claude or Grok?

#128

> I dropped eleven LLMs into a 2D battle royale and made them play 30 games. One won 43% of the matches. Three never won a single game. The cheapest model in the lineup beat the most expensive one by 27x on cost per win. Please learn how to write with AI without giving away that it was written by AI.

What about that makes you think it was written by AI?

Since you asked...I've gone to the effort to pull out the parts of the article that I think show it:

"That’s the part most benchmarks can’t see, and it’s what this post is about." Classic "it's not x, it's x", shows up in various forms throughout the article.

"To me, this is the most fascinating finding from this entire experiment - we saw very clear alignment tax being paid by certain models, which directly impacted their performance in this zero-sum game." - Usage of em dash. Now, yes, there's nothing wrong with using em dashes. But this feels like a weird place to use one. Also I counted at least 6 other emdashes in this article. Most people do not use em dashes that often.

"and a memory system that kept doubling down on what worked without second-guessing or doubting itself." - Doubling down is a classic Claudism.

"I want to be careful here..." - "wanting to be careful here" is another classic Claudism.

"The same game world, completely different results when in a different “task”." - "same X, completely different X" is another common one from Claude, as proofed by the repeated pattern later down: "These models were all given the same rules, same game world, and same tools, but each of them approached the game on a personality-level that is completely different from each other."

"It begs the question" - author used this twice in the article.

I'm guessing the author wrote a draft and then had Claude spruce it up a lot. I could be wrong and I'd be happy to be proven otherwise.

Re: A robot is sprinting towards you. Do you want it running on Claude or Grok?

#129
post #8

If the robot appears to be bringing me a taco, it would probably penetrate all of my defenses. Grok is currently more likely than Claude to arrive with the taco without being stopped by an export control directive.

I'm reminded of the Alameda Weehawken burrito tunnel: https://idlewords.com/2007/04/the_alameda_weehawken_burrito_...

The single most implausible idea in that article is that New York City would be able to so completely outbid the SF Bay Area for burritos.

Re: A robot is sprinting towards you. Do you want it running on Claude or Grok?

#130
post #17

Earlier quoted context omitted.

What if the car can talk you through the medical procedure?

How many times have you been to a hospital and thought, I could have fixed that myself if only I'd known how? With no equipment. In my case, never.

That article was way longer than I thought it would be.

https://en.wikipedia.org/wiki/Self-surgery

Post reply on HN