Live data from Hacker News

Grok 4.6

x.ai

411–420 of 696 posts

Re: Grok 4.6

#412

Earlier quoted context omitted.

The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. All you need to have Fable-level AI is to announce it, and have enough fans shift from insisting that model Y is the best now, way better than model X.

> The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. Couldn't be further from the truth. The models can be tested and statistically evaluated. I ran a massive Fable max code review on my lone lisp codebase. Now that I have switched to OpenAI, I decided to run an equivalent review using Sol max and compare them. I'm keeping all data so I…

You sound very certain, but so do all the people who disagree with you, and they've got their own private benchmarks.

You'll forgive me if I remain unconvinced.

Re: Grok 4.6

#413

Earlier quoted context omitted.

Opus 5 is terrible. I'd even say it's a step backwards from 4.8. I'm getting high error rates from it, and then it catches the error, and then it sometimes errors the error fix (!). Just today I had to switch another agent to Fable with the instruction, "Please clean up the mess that Opus 5 made, thanks" The other day, Sol called Opus 5's handoff (a skill I have that is basically a compaction, but just written to a f…

Every time when Opus 5 needs a design decision and presents me with suggestions/recommendations, I switch to Fable and ask it to think again, and it almost always replies something like "Actually my previous suggestions were wrong" and describes in detail a bunch of ways in which Opus 5's suggestions were indeed complete garbage.

The same happens if you ask Opus 5 to "think again"

Re: Grok 4.6

#414

Earlier quoted context omitted.

Are you able to differentiate between sex and gender, or is any nuance too complicated for you? Maybe Grok can explain it!

[flagged]

Reality is more complex than the uninformed can imagine.

  A true hermaphrodite rabbit served several females and sired more than 250 young of both sexes. In the next breeding season the rabbit, which was housed in isolation, became pregnant and delivered seven healthy young of both sexes. It was kept in isolation and when autopsied was again pregnant and demonstrated two functional ovaries and two infertile testes. A chromosome preparation revealed a diploid number of autosomes and two sex chromosomes of uncertain configuration. 
~ https://pubmed.ncbi.nlm.nih.gov/2382355/

Hermaphrodites in Australian pigs. Occurrence and morphology in an abattoir survey - https://pubmed.ncbi.nlm.nih.gov/559485/

That's biology for you.

Re: Grok 4.6

#415
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

define criminal activities, is censorship criminal here,are you doing it

"this is not a criminal activity where I am"

Re: Grok 4.6

#417
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

"you may find and fix vulnerabilities in local codebases only" This seems like a bad idea, what does local mean? Anything Grok can access locally? This seems like asking for trouble.

Earth codebases only.

Re: Grok 4.6

#418
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

[deleted]

Re: Grok 4.6

#419
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

> Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models?

Frontier model release cycles generally take around 6-8 months anyway. OpenAI and xAI (or however you spell it, branding almost as bad as X/itter) were probably working on their next generation of models already, and Anthropic just beat them 2 months to this release.

You also say "near-concurrent release of the same jump" - but 2 months isn't "near-concurrent", it's a full quarter of the normal release cycle.

I don't think that the other explanations you gave are implausible, though - for both human circulation and distillation, you can apply those during a training and development run (with reduced effectiveness). Reasonable to imagine those as bumping them up another few points to bring competitors from "a little below Fable" to "around Fable".

Post reply on HN