Live data from Hacker News

Grok 4.6

x.ai

431–440 of 696 posts

Re: Grok 4.6

#431

It's crazy that I'd literally trust a Chinese AI company with my data over anything Musk is involved with. Like, even if you don't care about (or even like) his politics and can look past how unlikable he comes off as, the damage he's done to his own reputation in this domain just makes using his products like this a no-go. He's literally so rich that he can get caught personally looking through chat sessions and it…

> He's literally so rich that he can get caught personally looking through chat sessions and it wouldn't slow him down a bit. Looking through chat histories is boring, mundane stuff. He's richer than that, think bigger. I think he could kill a random person in front of thousands, and by the next day we'd see articles arguing why the random person actually deserved it and why it's not that bad. Whatever consequences w…

> I think he could kill a random person in front of thousands, and by the next day we'd see articles arguing why the random person actually deserved it and why it's not that bad

Think even bigger. How many deaths is he responsible for as a result of DOGE cuts to overseas aid? This seems to be water that passed under the bridge a long while ago as far as 'societies attention' goes.

https://hsph.harvard.edu/news/usaid-shutdown-has-led-to-hund...

https://www.doge-impact.org/

Re: Grok 4.6

#432
post #375

I am noting Opus 5 is omitted. Interesting as I thought it benched better than Fable 5 in a few benchmarks.

The Opus 5 release was a perfect example of how useless these benchmarks are for a head to head model comparison. Anthropic published a post showing Opus 5 beating Fable in almost every eval but then added a disclaimer that it was still a tier below Fable in intelligence (and thus pricing). So then what did all the numbers represent exactly?

My experience has been a difference between "applied intelligence" and "breadth of intelligence".

Fable is the theoretical computer scientist while Opus is the Staff engineer who will implement it.

I find that Opus has continually done better on tasks mechanically but if it misunderstands even one thing -- it might waste your time doing the wrong task well.

I've found Fable to be the better thinker, filling it the gaps in your spec, and having a common sense understanding of what you likely meant.

Re: Grok 4.6

#433

[flagged]

> I’m seriously considering switching to Grok What's holding you back? According to your post history you've been calling Grok "awesome" for months now: https://news.ycombinator.com/item?id=47988753 Is there any part of Anthropic's offerings that you're struggling to leave behind?

Busted! lol

Re: Grok 4.6

#434
post #379

Earlier quoted context omitted.

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

What do you think a human brain is…

an animal organ?

Re: Grok 4.6

#435
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It's about chips with a large enough scale up domain. Larger domain allows for bigger model, which is what's driving this jump. You've got to get the chips, test them, tune kernels, then start a big pre train, mid & post-train, and only then do you actually get the model. So it takes time. Anthropic got there first partly because they use different hardware (TPU I think, maybe Trainium) which had larger scale ups earlier.

Re: Grok 4.6

#436

Earlier quoted context omitted.

The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

For a kitchen knife this was okay, but the AI firms think that they’ve built a drone that’s the size of a phone but can fly 100km and can hold a kitchen knife. It might be used to assassinate someone before others can react or even catch them.

Re: Grok 4.6

#437

Earlier quoted context omitted.

The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

[flagged]

Re: Grok 4.6

#438

Earlier quoted context omitted.

This is like saying the person you see in the mirror is categorically a human being because both of you produce similar reflections of light rays

The person I see in a mirror is a human being. The person I see in a mirror is me. What do you think a mirror is?

Then why is he left-handed?

Re: Grok 4.6

#439

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

Criminal activity by which countries laws?

Is this an attempted gotcha?
Post reply on HN