Live data from Hacker News

Grok 4.5

x.ai

601–610 of 1001 posts

Re: Grok 4.5

#601

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

You're not voting it into office (yet, anyway. Haha). Aren't we all just using these models to write code?

Re: Grok 4.5

#602

Earlier quoted context omitted.

Studies on political bias in models consistently show that LLMs lean politically left. The only outlier is grok which leans right but by a smaller factor, according to this study for instance: https://arxiv.org/abs/2603.23841 Edit: adding some other studies that are easily retrievable with a quick search for those unsatisfied with the first one - https://arxiv.org/abs/2606.12922 https://arxiv.org/abs/2412.16746 Claim…

[flagged]

People on the right say the same about the left

Re: Grok 4.5

#603
post #5

It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 . I guess the Cursor data was very useful.

I have a theory that xAI has one of the largest clusters but with far less traffic + tokens to process bc its less popular than its competition, and xAI can pass the savings on to the end user.

No need for theorizing. xai is selling their excess capacity to Anthropic and Google with large markup

Re: Grok 4.5

#604
post #395

Earlier quoted context omitted.

Cursor has had a good AI product tons of people used for real work for 2yrs (up until recently when the Claude gap widened significantly) while Microsoft/Github has just been pretending they do with Copilot and awful Github AI integrations nobody likes. Meanwhile Github's code has already been vacuumed up by all the models by now.

You would be surprised how many enterprises are allowed to use only Copilot for all tasks due to Microsoft deals.

yeah, but that's due to enterprise commitments that MS won't train on the user interactions

Re: Grok 4.5

#605

(from Cursor's blog) > Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments. This is what the big money was for. Cursor is the first big player that had rea…

> You use the previous gen model to prepare datasets for the next model iteration I've read multiple times that this approach is harmful in training. You're essentially describing what many call distillation, but it's only useful in post training to guide behavior, it teaches how to behave, not how to think. I might be wrong though and would be glad if someone more knowledgeable provided more insights.

There was one highly discussed paper. And about a month after publication much much stronger models were released. The models got better because they used such synthetic data.

Re: Grok 4.5

#606

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

from my experience, grok seems to be the least censored models among all big vendors, not sure if it's got injected any political bias (I know there was a rumor that it was primed with Elon's personal feed at one point in the past) but that alone makes it better than ChatGPT and Gemini.

On top of the model, Grok seems to always do many web searches for every prompt I throw at it, which makes better than even Gemini as a search engine replacement (you'd think Google must have nailed that usecase but nope). ChatGPT is too lazy in this regard and half the times just split out an answer right away.

Re: Grok 4.5

#607
post #92

First impressions: - Very fast, easily beats GPT 5.5/Opus 4.8/GLM 5.2 because of higher t/s (around 90?) and very high token efficiency - Very good price, no contest vs GPT and Opus which are very overpriced if you pay API costs, and probably cheaper than GLM 5.2 when you take into account the token efficiency. - Will take quite a while to get a feel for how smart it is, but it's definitely good, I'd say in the same…

The input price is quite high. That's what gets you.

Re: Grok 4.5

#608

I am amazed at people's willingness to use Grok. The company is so transparently morally bankrupt. They're the only AI company that seems okay with CSAM (or at least don't do as much to stop it) Why give them money? It would be one thing if they were the only game in town but thats definitely not the case.

> okay with CSAM

based on what? I tried and it could not even generate adult nudity. Earlier Grok Image was poorly censored but now they have their filters in place.

Re: Grok 4.5

#609

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

What a crazy thing to say explicitly about the model that _avoids_ that kind of stuff. Did we already collectively forget about the ChatGPT "misgendering is worse than thermonuclear war" case, or Gemini picturing the Founding Fathers as black?

Anything political, I always go to Grok first. It's the only one that has bled for not just trying to play it super safe with political correctness, but trying to be impartial.

Re: Grok 4.5

#610

I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?

Grok has the most balanced answers according to many tests. I think you should try and ask gemini and claude some political questions.
Post reply on HN