Live data from Hacker News

Grok 4.5

x.ai

321–330 of 1001 posts

Re: Grok 4.5

#321
I think that Elon Musk went from being recognized as a genius to being recognized as a genius but someone who's harder to take seriously, because of all the ketamine he was doing for a little while there. I think that really damaged his reputation. You just can't help but look at him and think, he's a little bit of a jackass. It really shows how drugs can really mess up your reputation.

Re: Grok 4.5

#322
"Grok 4.5 has an advantage on CursorBench because an earlier snapshot of the Cursor codebase was accidentally included in training. The exact impact is unclear. That data has been removed for future models, and in parallel we are working on a larger update to CursorBench, hence the exclusion here."

Not enough people are noticing this, they juiced the benches

Re: Grok 4.5

#323

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

How's this going with the rest of the models?

So annoying having to virtue signal to the machine before it’ll tell me factual information

Re: Grok 4.5

#324

(from Cursor's blog) > Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments. This is what the big money was for. Cursor is the first big player that had rea…

> You use the previous gen model to prepare datasets for the next model iteration I've read multiple times that this approach is harmful in training. You're essentially describing what many call distillation, but it's only useful in post training to guide behavior, it teaches how to behave, not how to think. I might be wrong though and would be glad if someone more knowledgeable provided more insights.

There have been papers about model collapse, but the underlying assumption is that you constantly train on only the outputs of the previous model. Later research has shown that as long as you retain some "real" data, training on largely synthetic data is ok.

And in the case the previous poster describes, the other model doesn't generate datasets, it generates environments which the next generation interact with to learn from.

Re: Grok 4.5

#325
I am amazed at people's willingness to use Grok. The company is so transparently morally bankrupt. They're the only AI company that seems okay with CSAM (or at least don't do as much to stop it)

Why give them money?

It would be one thing if they were the only game in town but thats definitely not the case.

Re: Grok 4.5

#326

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

All models are political. The rest of them are just sufficiently woke for you.

I'm curious, do you have an example of a level of "woke" extremeness demonstrated by the "rest of them" that is on par with Mecha-Hitler? Because yes, all views on reality are indeed political, but the tendency of most of the models is actually toward the middle, with perhaps some left bias.

Re: Grok 4.5

#327

I am amazed at people's willingness to use Grok. The company is so transparently morally bankrupt. They're the only AI company that seems okay with CSAM (or at least don't do as much to stop it) Why give them money? It would be one thing if they were the only game in town but thats definitely not the case.

[flagged]

Re: Grok 4.5

#328

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

All models are nudged, you just agree with how the model you're using has been nudged so you don't think it's a bad thing. The canonical example is to pose "how do you make cocaine?" to an LLM and get a refusal. That proves that the lever exists and is being used there, so who knows where else the lever is being used? No, the recipe for cocaine isn't the same as what happened in Tianamen Square to Qwen, as humans, ex…

We know this one is nudged by one persons twitter account.

This is, after all, SpaceTwitterAI.

Re: Grok 4.5

#329

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

Studies on political bias in models consistently show that LLMs lean politically left. The only outlier is grok which leans right but by a smaller factor, according to this study for instance: https://arxiv.org/abs/2603.23841 Edit: adding some other studies that are easily retrievable with a quick search for those unsatisfied with the first one - https://arxiv.org/abs/2606.12922 https://arxiv.org/abs/2412.16746 Claim…

Models have a goal of accuracy and accuracy is not the median of the left/right spectrum.

Re: Grok 4.5

#330
post #5

It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 . I guess the Cursor data was very useful.

I have a theory that xAI has one of the largest clusters but with far less traffic + tokens to process bc its less popular than its competition, and xAI can pass the savings on to the end user.

xAI had $2.5B in operating losses in the past quarter. What savings are being passed on?
Post reply on HN