Live data from Hacker News

Grok 4.5

x.ai

211–220 of 1001 posts

Re: Grok 4.5

#211
post #3

How popular is Grok compared to other companies models for SWE tasks? I almost never hear it talked about against OpenAI's or Anthropic's products

They had two big substantive flaws on top of the political stuff. Aside from a brief window last summer Grok has been behind the curve for coding, and before the Cursor acquisition they didn’t have a harness. Now they have an Opus tier model and a real harness they have at a minimum the opportunity to undercut the competition on price. And with the 5T and 10T models being trained on Colossus 2 they have the possibility to leap ahead.

Re: Grok 4.5

#212

Earlier quoted context omitted.

Sonnet 5 is a huge token hog, though, it uses far more reasoning tokens than Opus models while being priced at $2/$10 with promo, and $3/$15 (usual Sonnet price) afterwards.

I'll probably get hate for it, but I was not impressed by Fable, I felt like it was just Opus with more tokens for thinking. I feel like the second I turned on Fable I drained my usage more quickly, despite them billing it as though it were Opus level of usage. The value is just not there for me. I wish they could make Haiku remain low-cost and drastically more capable to the point you could use only Haiku.

Fable needs more... ambitious tasks than Opus to tell the difference and let me tell you the difference is there.

Simple tasks are simply saturated just like simple benchmarks. There's a level of intelligence where you simply don't need more for some things.

Re: Grok 4.5

#213
post #207

Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.

SpaceX needs to keep raising many billions every year. The rockets part isn't going to make money for a long time, so diversion tactics https://news.ycombinator.com/item?id=48828648 Also Elon has a grudge with Sam Altman and wants to beat him

even more so after losing the lawsuit, imo

Re: Grok 4.5

#214
post #203

Very hard for me to imagine this getting beyond a low-single-digit market share. I don't understand the strategy of xAI burning money on this.

I think the strategy is pretty obvious

Re: Grok 4.5

#215
post #5

It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 . I guess the Cursor data was very useful.

I have a theory that xAI has one of the largest clusters but with far less traffic + tokens to process bc its less popular than its competition, and xAI can pass the savings on to the end user.

Re: Grok 4.5

#216

Do we have any proof that this was made by xAI and isn't some Chinese open model running with modifications? Their inital image generation was a wrapper around Flux.

Even if they did start from an open model base, does (or should) that matter if it performs well?

Genuinely asking.

Re: Grok 4.5

#217

Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.

Grok is the #1 uncensored easily-available model, and it's also tightly integrated with Twitter.

Great if you want to make virtual child porn, I guess.

Re: Grok 4.5

#218
post #179

Earlier quoted context omitted.

Ollama is just a local app wrapper/cloud service serving third party apis and models idk why it made it into this list tbh

Claude isn’t a model either.

He probably asked AI to make the list for him

Re: Grok 4.5

#220
post #29

Earlier quoted context omitted.

Grok Build sucks compare to composer 2.5. Just use compose 2.5 and you'll have basically unlimited usage on the 40$ plan.

It is hard to evaluate the model performance of Composer 2.5 when Cursor's harness is so awful compared to the others on the market.

In what way? I spend more of my time managing than hands on lately so I legitimately don’t know.
Post reply on HN