Live data from Hacker News

Grok 4.5

x.ai

21–30 of 1001 posts

Re: Grok 4.5

#21
post #3

How popular is Grok compared to other companies models for SWE tasks? I almost never hear it talked about against OpenAI's or Anthropic's products

Because of the of the political stuff, they have a bad reputation I think and are taken less seriously (I feel this way). They have an opportunity imo to break free from that and just not do the gatekeeping / condescension that the other providers are starting, and become more mainstream.

Re: Grok 4.5

#22
post #3

How popular is Grok compared to other companies models for SWE tasks? I almost never hear it talked about against OpenAI's or Anthropic's products

[deleted]

Re: Grok 4.5

#24
post #8

Of the 3 models I tried, Grok did the best at making an iOS app I wanted for personal use (a bike computer with specific qualities). (Claude just gave up and did an HTML/CSS implementation but I insisted on native SwiftUI+Metal.) Grok definitely fumbles sometimes, but I have been surprised what it CAN intuit versus me having to micromanage it. (I am not an iOS developer, so getting something specific that I needed in…

Was this in Claude Code for Claude? Did you use a weaker model like Haiku? Claude should absolutely not be as bad as you said.

I tried Claude Code with XCode once, I already use CC exclusively, either in the CLI or with Zed (mostly CLI now), and it was pretty unstable. I wish Apple would QA their products more. It seems to me the best way to use Claude Code for anything is stand-alone.

Re: Grok 4.5

#25
post #3

How popular is Grok compared to other companies models for SWE tasks? I almost never hear it talked about against OpenAI's or Anthropic's products

They were missing a harness like Claude Code or Codex (terminal). However they recently released Grok Build, which is probably the fasted I've used, in terms of responsiveness, but didn't have a model at Opus 4.7/8 level. The thing is if they add 4.5 to Grok Build and keep improving the harness I think it can compete (cheaper and faster).

Re: Grok 4.5

#26
post #16

Not available for Europeans yet. :(

I think it should be available through Cursor? EDIT: Tested myself, it's actually NOT available from EU. But with a Swiss VPN it works :)

We will probably see it when it's available for everyone.

This is the first time I see a lab region locking a model though.

Re: Grok 4.5

#27

Is there a reason the AI companies usually announce new products so close to each other. Like not just the same day but literally hours apart. GPT Live then an hour later Grok 4.5. As if they try to one up. I expect something new from Anhtropic as well today.

I think this one is just a coincidence, bound to happen given the pace of releases

For exact timing, probably 10-11am Pacific is just optimal for normal working hours

Re: Grok 4.5

#28
Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years?

Re: Grok 4.5

#29
post #5

It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 . I guess the Cursor data was very useful.

Now if they could have an "equivalent" to Claude's $100 plan with similar compute limits. I have the $40 a month version of Grok and I get a max of like 8 hours of "non-stop" Grok Build coding, per month.

Grok Build sucks compare to composer 2.5. Just use compose 2.5 and you'll have basically unlimited usage on the 40$ plan.

Re: Grok 4.5

#30
post #21
post #3

How popular is Grok compared to other companies models for SWE tasks? I almost never hear it talked about against OpenAI's or Anthropic's products

Because of the of the political stuff, they have a bad reputation I think and are taken less seriously (I feel this way). They have an opportunity imo to break free from that and just not do the gatekeeping / condescension that the other providers are starting, and become more mainstream.

[flagged]
Post reply on HN