Live data from Hacker News

Grok 4.6

x.ai

421–430 of 696 posts

Re: Grok 4.6

#421

Earlier quoted context omitted.

> The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. Couldn't be further from the truth. The models can be tested and statistically evaluated. I ran a massive Fable max code review on my lone lisp codebase. Now that I have switched to OpenAI, I decided to run an equivalent review using Sol max and compare them. I'm keeping all data so I…

You sound very certain, but so do all the people who disagree with you, and they've got their own private benchmarks. You'll forgive me if I remain unconvinced.

If I sounded certain, it was not intentional. I made sure to hedge my statistical claims with "seems" and "looks like". I'm no AI lab, I'm just a random subscription user trying to get the most value out of them.

I'm just saying it's not wise to simply put all these models in the same bucket and say any differences are due to vibes or hype. They are clearly different. We can and should scrutinize the testing methodology but it's not exactly fair to just ignore the results.

I don't intend for my benchmark to be private. The core component of my test is my parallel code review skill which is already on my GitHub. I'll be publishing the results on my website when it's done. Anyone could take the skill and reproduce the test using multiple models against any codebase out there, then analyse the depth of each model's findings.

Re: Grok 4.6

#422

Earlier quoted context omitted.

[flagged]

Reality is more complex than the uninformed can imagine. A true hermaphrodite rabbit served several females and sired more than 250 young of both sexes. In the next breeding season the rabbit, which was housed in isolation, became pregnant and delivered seven healthy young of both sexes. It was kept in isolation and when autopsied was again pregnant and demonstrated two functional ovaries and two infertile testes. A…

>Reality is more complex than the uninformed can imagine.

Reality is more simple than the delusional can image. Yes, intersex is weird. Literally states "hermaphrodites" in the title.

Glad you agree males cannot be pregnant.

Re: Grok 4.6

#423
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

To be really clear, Fable is Opus.

Anthropic finished a new pre-training run, Opus-sized models got enough of a jump they could have released Fable as Opus 5... but the economics of Opus models weren't where they wanted.

Being the masters of distribution that they are, instead of announcing a massive price hike, they just introduced a new tier and promoted Sonnet-sized models to Opus.

That's why every Opus after 4.6 has had such mixed feedback: smaller model with more RL can only make up so much ground, especially on vibes (which are hard-to-impossible to build a reward for)

(I mention all of this because if they'd just released Opus 5, no one would be asking "why is it a few months later everyone caught up to the latest release"... that's always how it works)

Re: Grok 4.6

#424

Earlier quoted context omitted.

I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…

The US govt trusts SpaceXAI for defense and high security missions. The idea they are lying about contracted AI services is absurd. They're also a public company which beings even more oversight than openai / anthropic.

Is this top tier satire or pure naivety? You're arguing that military contractors or public companies couldn't possibly be corrupt or nefarious?

Re: Grok 4.6

#425
post #417

Earlier quoted context omitted.

"you may find and fix vulnerabilities in local codebases only" This seems like a bad idea, what does local mean? Anything Grok can access locally? This seems like asking for trouble.

Earth codebases only.

[deleted]

Re: Grok 4.6

#426

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...

> The alternative is Claude-style "safeguards" aka censorship

Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

Re: Grok 4.6

#427
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

define criminal activities, is censorship criminal here,are you doing it

Censorship being explicitly a function of government, at least in legal terms, seems pertinent here.

Re: Grok 4.6

#428

>Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. As a designer, I'm always hesitant to believe these statements until there's independent comparisons between the old & new model, as well as comparisons to human made flows. Design can be so…

A relatively frequent pattern for AI announcements is "it's now really good at X", almost always said by someone who is not an expert in X.

Marketing Senior: We need someone to quote as saying "it's now really good at X", ideally someone who is an expert in X.

Marketing Junior: We've approached as many experts in X as we could, and demonstrated the new X capabilities to them, and no one wanted to be quoted by name saying that phrase, or even slightly watered down versions of that phrase.

Marketing Senior: How many celebrities do we have contact details for?

Re: Grok 4.6

#429

Fable-like intelligence, beats GPT-5.6-Sol on most benchmarks, cheaper than Kimi K3 on API and quite generous usage on Cursor subscription.

I like Grok, but I don't think that it's quite Fable-tier. It's good, but I think the position that it occupies on the Pareto frontier is a little more toward the "cheap" side and a little less toward the "intelligence" side.

Re: Grok 4.6

#430
post #29
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

4) There's nothing terribly special about Anthropic. No moat.

the "no moat" stuff is dumb. There are only like 4 companies in the world that have cutting edge LLMs so there is definitely a moat.
Post reply on HN