Every time I get excited about Grok’s performance on benchmarks and demo videos, I test it myself and end up disappointed. I'll give this one a try with a grain of salt and lowering my levels of expectations
[flagged]
Grok 4.5
381–390 of 1001 posts
Re: Grok 4.5
#382Earlier quoted context omitted.
Studies on political bias in models consistently show that LLMs lean politically left. The only outlier is grok which leans right but by a smaller factor, according to this study for instance: https://arxiv.org/abs/2603.23841 Edit: adding some other studies that are easily retrievable with a quick search for those unsatisfied with the first one - https://arxiv.org/abs/2606.12922 https://arxiv.org/abs/2412.16746 Claim…
Models are tuned to give ethical responses and right-leaning responses are judged by RLHF process to be unethical. That's your problem. It's not 'left' or 'right' to be ethical, but if one side is inherently antisocial and unethical then it's going to naturally create an appearance of bias toward the other.
Re: Grok 4.5
#383It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50. And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049 . I guess the Cursor data was very useful.
Grok is stuck in a difficult place - not the best model at anything, and not the cheapest either. It's hard to make a case for using it on any dimension, even before you factor in the history (I'm not sure suggesting the company uses the model that refers to itself as "MechaHitler" is the way to a promotion).
Re: Grok 4.5
#384I am amazed at people's willingness to use Grok. The company is so transparently morally bankrupt. They're the only AI company that seems okay with CSAM (or at least don't do as much to stop it) Why give them money? It would be one thing if they were the only game in town but thats definitely not the case.
[flagged]
Re: Grok 4.5
#385Earlier quoted context omitted.
Studies on political bias in models consistently show that LLMs lean politically left. The only outlier is grok which leans right but by a smaller factor, according to this study for instance: https://arxiv.org/abs/2603.23841 Edit: adding some other studies that are easily retrievable with a quick search for those unsatisfied with the first one - https://arxiv.org/abs/2606.12922 https://arxiv.org/abs/2412.16746 Claim…
Reality leans left in many respects, principally the non-economic ones. It's a simple consequence of the same trend of overall social and educational progress that allowed these models to be developed in the first place. There is a reason they came out of San Francisco, and not Russia, Iran, or Oklahoma. To get a right-biased response from an LLM, you have to deliberately bias it... which is exactly what Musk did. Ne…
Re: Grok 4.5
#386Earlier quoted context omitted.
Studies on political bias in models consistently show that LLMs lean politically left. The only outlier is grok which leans right but by a smaller factor, according to this study for instance: https://arxiv.org/abs/2603.23841 Edit: adding some other studies that are easily retrievable with a quick search for those unsatisfied with the first one - https://arxiv.org/abs/2606.12922 https://arxiv.org/abs/2412.16746 Claim…
a non-peer-reviewed preprint with an high schooler as the first author is the best citation you can come up with?
Re: Grok 4.5
#387Re: Grok 4.5
#388The anti-Musk stuff would qualify as brigading in nearly any other community. It shocks me that people have such a visceral, irrational engagement with anything in Musk's orbit. I probably shouldn't have, but I expected better from the HN crowd for some reason. It's an excellent model. GPT 5.4/5.5 level, some things better, others not, but extremely fast. A wonderful technical improvement. If a Chinese company or ran…
Re: Grok 4.5
#389Earlier quoted context omitted.
How's this going with the rest of the models?
So annoying having to virtue signal to the machine before it’ll tell me factual information
Honestly though, that pales into comparison with the fable censorship. I never realized how many metaphors I use are either biological or security related in nature (ex: asking claude to reverse engineer something, in the metaphorical sense of the word). And the best part is I can't even tell the fable instance "you can't talk about mitochondria or you'll die" because then he'll go "of course I can, this is a legitimate scientific topic. The mitochondria is the power-BLAM [slumps over dead, Opus 4.8 crawls over his dead body and starts gaslighting me]"
Re: Grok 4.5
#390Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.
Previously they were a distant fourth. They're not going to single-shot catch up to OpenAI or Anthropic, but they moved up the ladder one rung. In the short term labs are not profitable, although supposedly Anthropic is close. But Amazon was also famously unprofitable for many many years, and then won huge. Current profits or lack thereof are not necessarily important to investors: what's important is they believe in…
Amazon was 'unit profitable' very early.
Yes - it's not unreasonable for Elon to bet long horizon ... there are after all many car companies, why not AI?
He's already winning gov. contracts, that could continue.
It's an odd bet but not entirely wrong or dubious.