Live data from Hacker News

Grok 4.5

x.ai

591–600 of 1001 posts

Re: Grok 4.5

#591
post #29

Earlier quoted context omitted.

Grok Build sucks compare to composer 2.5. Just use compose 2.5 and you'll have basically unlimited usage on the 40$ plan.

Every time I use Composer 2.5 I have to spend a bunch of time cleaning up its mistakes. It is unusable compared to GPT 5.4 or 5.5. My time is more valuable that I will use a model that doesn’t f** up my code base.

Isn't Composer 2.5 designed to be used from the Cursor harness, and is otherwise not that useful?

Re: Grok 4.5

#592
grok 4.5 managed to debug and fix and issue that caused an incident for my project yesterday. I ran a multiagent debugging session first with grok 4.5 high, then it found the root cause and implemented a small fix in k8s manifests, deployed and verified the fix, all in under 30 minutes. the day before it took me 3+ hours of debugging and poking around in several sonnet 5 medium sessions to at least figure out what was going on - and I didn't. in terms of context usage, grok used ~115.9K for the whole session.

Re: Grok 4.5

#593

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

Unfortunately all of the major AI model providers are massively incentivized to fit their models to various political narratives, especially through historical denialism. The "diverse 1940s German soldiers" debacle from Google comes to mind, or perhaps "nothing of note happened at Tiananmen Square" from any of the Chinese models.

Do you genuinely believe those two examples are comparable? Image generation and recitation of historical consensus are two very different domains, primarily because of how much more information dense an image is than a blurb of text.

Put more lightly, if I ask a model to “generate an image of a soccer player”, what’s the most politically neutral option of the following:

  - Make them white, because of American cultural hegemony
  - Make them brown, because that’s a more globally average skin tone
  - Try to infer the user’s skin tone based on personal and location data, and use that for the player
  - Try to infer the user’s gender based on personal data, and use that for the player
  - (*) Browse the news for the most famous or trending soccer player right now, and use that player in the image
  - Do the same as the previous step, but make it more local to the user
  - Use the data encoded within the LLM to infer what the most likely appearance of a soccer player would be, which is then of course biased by what your data is and how it was collected
  - etc. etc. etc.
IMO there’s no option that won’t piss someone off, because I’m sure the knee jerk reaction is to choose the one I indicated with a (*), but now if you do that with the prompt “generate an image of a ketamine addict” or “generate an image of a serial adulterer” you may get into some trouble.

There is no neutral option, so if you’re either genuinely upset, or feigning being upset in order to virtue signal, it’s not that there’s an objective alternative that you prefer, it’s that you’re upset because it doesn’t match your subjective preference.

Re: Grok 4.5

#594

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

People, even in tech, even in AI, are allowed to have opinions that differ from yours. Cancellation doesn't work anymore: you can't get away with presenting your side as normal and other side as deviant and let ostracism win your idea's battles without your idea having to fight for itself.

The reason it's cringe is that you can really only understand an idea by placing it in tension with other ideas. Remove your idea's competition and you remove an incentive to explore the weaknesses of your own. Ideological protectionism, just like the economic kind, breeds weakness. Your idea mutates and, without fitness feedback, drifts into a ridiculous parody of itself. Your idea ceases to be a thing that can live on its own and comes to depend on the protectionism for it's survival. Yet, the more ridiculous your idea becomes, the more protectionism it needs to compensate. One day, your idea collides with other ideas (which are still out there) despite your best efforts to shield it attack, and when it does, you're shocked by how weak your coddled, mutated idea really is compared to the original form one might remember.

All the people out there criticizing Grok, or Grokipedia, or whatever for espousing the wrong ideas are ultimately undermining their own. Even if you don't believe in high-minded mumbo-jumbo about the value of free speech, even if you just want your side to win, trying to shame models like Grok into not existing is foolish and undermines your goals.

Re: Grok 4.5

#595

Earlier quoted context omitted.

Grok is the #1 uncensored easily-available model, and it's also tightly integrated with Twitter.

Is uncensored a selling point? What do people use uncensored Grok for (like, real use cases) that they can't or won't use other LLMs for? Literally the only thing I can think of is generating bad porn of unconsenting people.

> What do people use uncensored Grok for (like, real use cases) that they can't or won't use other LLMs for? Literally the only thing I can think of is generating bad porn of unconsenting people

untrue. There's a full thread about it: https://news.ycombinator.com/item?id=48837162 - but as much as I love Claude products, nothing's more aggravating than it refusing to help me diagnose a stack trace because it "violates Anthopic policy".

Re: Grok 4.5

#597

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

I don't care about the politics, I wouldn't trust anything made by Musk.

Re: Grok 4.5

#598

Earlier quoted context omitted.

I do a lot of native iOS development using Opus 4.8 (and I used 4.7/4.6 before this). I have a very hard time with this comment, were you using Opus or something else?

Same. A few months ago I pointed Opus 4.6 at a mid-size Vue app and told it to create the iOS equivalent using SwiftUI, and it nailed it. I broke the process down to phases and reviewed each phase, but within about ten days I had a functioning iOS app that had full feature parity.

I've done the same with DS-V4-Pro, GLM-5.2, MiMo-2.5-Pro, etc. - this is a task pretty much any agentic model can handle nowadays.

(I do the reverse currently where I implement a macOS front end natively, and then just let the agent rip on porting that to an HTTP API server + React/TypeScript/HTML/CSS frontend, because it's significantly easier to have agent loops fiddle with making macOS apps than it is to fiddle with a web browser and CSS.)

Re: Grok 4.5

#599

Announcement from Cursor, whose team also trained the model: https://cursor.com/blog/grok-4-5 . Notably: > Grok 4.5 and Composer 2.5 are two different model weight classes, and we're excited to support both sizes and weights. Composer 2.5 will remain offered, and we will release new models of this size going forward.

Composer 2.5 is 1T total/32B active (based on Kimi 2.5), while Elon publicly said Grok 4.5 is 1.5T parameters total. Hardly a different weight class. The API cost difference is ~2.5x, probably because xAI has much higher costs to recoup.

I would be utterly shocked if Grok 4.5 only has 32B active, given the results I am seeing from it. My guess is it's somewhere around 90B-100B active.

Re: Grok 4.5

#600
post #406

Every time I get excited about Grok’s performance on benchmarks and demo videos, I test it myself and end up disappointed. I'll give this one a try with a grain of salt and lowering my levels of expectations

I am trying to benchmark it now, but: - It doesn't seem available in EU (?) - Using a VPN seems to sort of fix it, but it's way slower than I expected, when everyone was praising it, it feels like the speed is slowly ramping up - Cost is $2/$6 for https://aibenchy.com/compare/x-ai-grok-4-5-medium/z-ai-glm-5...

I concur that GLM-5.2 still seems "better" in my experimentation I did with it tonight, although Grok 4.5 is cheap if you can tolerate the way a Cursor subscription works.
Post reply on HN