Live data from Hacker News

Grok 4.5

x.ai

91–100 of 1001 posts

Re: Grok 4.5

#91
Props to them for including three benchmarks that actually seem to say something, instead of focusing on totally gamed benchmarks like regular SWE-Bench. That could mean this model is actually pretty close to the SOTA as the benchmarks indicate.

Most labs - including OpenAI and Anthropic, but also Google and Chinese labs - highlight their scores in benchmarks that have fixed, widely available answers. Those answers end up in the training data and so models can just regurgitate training data instead of actually doing the benchmark. As a result, most benchmarks often quoted are essentially meaningless for gauging model performance.

Terminal-Bench still publishes answers, but neither DeepSWE and SWE-Bench Pro do. Especially for DeepSWE it's been difficult for models to fake good results so far. SWE-Bench Pro does have weird outliers like good performance for e.g. the atrocious Muse Spark, but it also doesn't provide answers for the training data.

So either they're good, or they found a way to game DeepSWE. Given that the Cursor team previously published the well-received Composer 2.5 a good score here doesn't come out of nowhere, so this might hold up. Cursor has enormous amounts of training data to train good coding models with.

Re: Grok 4.5

#92
First impressions:

- Very fast, easily beats GPT 5.5/Opus 4.8/GLM 5.2 because of higher t/s (around 90?) and very high token efficiency

- Very good price, no contest vs GPT and Opus which are very overpriced if you pay API costs, and probably cheaper than GLM 5.2 when you take into account the token efficiency.

- Will take quite a while to get a feel for how smart it is, but it's definitely good, I'd say in the same tier as opus, occupying the lower end of that tier together with GLM 5.2.

Re: Grok 4.5

#93
post #75

With each release from the the other major labs, it becomes harder for Google to tell a compelling story about Gemini 3.5. Edit: Gemini 3.5 Pro . Expectations grow with each day it is not released.

[flagged]

Re: Grok 4.5

#94

Earlier quoted context omitted.

Competition. You don't want to lose your customers trying out the competitors updated and better product. Release on the same day and they won't be able to compare their new to your old.

But how do they know what day is that? Unless you have already something ready to be announced (and you just hold it until the very last moment, which doesn’t make sense, since you could just announce it asap)

It can also be ”we are done but wanna test it more and tweak it” and then ”oh they launched now. Let’s launch then as well”

Re: Grok 4.5

#95

Announcement from Cursor, whose team also trained the model: https://cursor.com/blog/grok-4-5 . Notably: > Grok 4.5 and Composer 2.5 are two different model weight classes, and we're excited to support both sizes and weights. Composer 2.5 will remain offered, and we will release new models of this size going forward.

Composer 2.5 is 1T total/32B active (based on Kimi 2.5), while Elon publicly said Grok 4.5 is 1.5T parameters total. Hardly a different weight class.

The API cost difference is ~2.5x, probably because xAI has much higher costs to recoup.

Re: Grok 4.5

#96
post #21

Earlier quoted context omitted.

Because of the of the political stuff, they have a bad reputation I think and are taken less seriously (I feel this way). They have an opportunity imo to break free from that and just not do the gatekeeping / condescension that the other providers are starting, and become more mainstream.

Even without the politics, Elon has shown that he will weaponize his platforms against people/companies he personally doesn't like (e.g. specific bans/demotions to external sites like Substack and Bluesky). Using Grok is therefore a supply chain risk and it's not nearly good enough to offset that risk.

I do just want to focus on the 'even without the politics' asterisk though because sometimes there is a risk people think everyone on x side (x meaning 'a given side', not x.com) is wrong

You can claim Elon bought x as some sort of power trip. Fine. Willing to entertain it, I have no dog in the fight. I'm not a member of the Elon fan club. And yet Twitter (under Dorsey though I don't think he was involved) was banning tons of people under guises of 'misinfo' that wasn't misinfo

Re: Grok 4.5

#98
post #63
post #28

Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years?

My gut feel is Anthropic is very technical and pedantic which makes their models really technical and pedantic. They're top at code and technical benchmarks but anecdotally I've found OpenAI to be significantly farther ahead for general usage. Opus 4.8 will burn 10k tokens trying to answer something 100% whereas GPT-5.5 will burn 2k getting it 90% which is good enough for many things. Some personal testing on a "help…

The problem is that the remaining 10% can bite you in bad ways.

I was in Cotswolds, UK a couple of months ago. For those of you who don't know, it's a rural region known for its "chocolate-box" villages and honey-colored limestone architecture. Basically, you go from village to village, most commonly via bus, taking in the sights and doing touristy stuff.

When planning the trip, my sister used ChatGPT, which helpfully (and relatively quickly) found the bus schedules and times for each hop.

Midway through the day, though, we ran into a huge problem: it turns out bus schedules are different on Sundays, and more limited. Which meant we couldn't actually go to our primary destination (the Model Village), and had to cut the trip short.

Yes, ChatGPT was quick and pleasant to use, but missed a crucial detail.

Afterwards I tried it with Opus and it did not make the same mistake.

Re: Grok 4.5

#99
post #29

Earlier quoted context omitted.

Now if they could have an "equivalent" to Claude's $100 plan with similar compute limits. I have the $40 a month version of Grok and I get a max of like 8 hours of "non-stop" Grok Build coding, per month.

Grok Build sucks compare to composer 2.5. Just use compose 2.5 and you'll have basically unlimited usage on the 40$ plan.

It is hard to evaluate the model performance of Composer 2.5 when Cursor's harness is so awful compared to the others on the market.

Re: Grok 4.5

#100

Announcement from Cursor, whose team also trained the model: https://cursor.com/blog/grok-4-5 . Notably: > Grok 4.5 and Composer 2.5 are two different model weight classes, and we're excited to support both sizes and weights. Composer 2.5 will remain offered, and we will release new models of this size going forward.

Composer 2.5 is 1T total/32B active (based on Kimi 2.5), while Elon publicly said Grok 4.5 is 1.5T parameters total. Hardly a different weight class. The API cost difference is ~2.5x, probably because xAI has much higher costs to recoup.

I could easily see Grok 4.5 being around 1:16 in terms of active parameters, so around 94B active parameters.
Post reply on HN