Live data from Hacker News

Grok 4.6

x.ai

311–320 of 696 posts

Re: Grok 4.6

#311
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Researchers moving between companies (and other ways that techniques get leaked) is the largest cause of this IMO. It's happening continuously, so I don't see why the timing makes it implausible. A really underrated strength of Silicon Valley is California's ban on non-competes that allows this to happen and ensures robust competition between model providers both for talent (increasing salaries for workers) and in the marketplace (reducing prices for consumers). If OpenAI had been located in New York instead then Anthropic could never have succeeded, for example.

But I think the other reason you didn't mention is the timing of new compute coming online. Compute is the major factor limiting the training of these models and new datacenter investments are bearing fruit at around the same time.

Re: Grok 4.6

#312
post #306
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It was said at the time that xAI acquiring Cursor was very smart because it would give them access to years of agent coding traces from millions of users. $60B in SpaceX stock for Cursor was a bargain Data + compute + being competent and smart enough to ship. fwiw I don't think these are yet Fable level - the difference tends to get discovered in the long tail of tasks - but they're close enough, they're cheap, and t…

> $60B in SpaceX stock for Cursor was a bargain

Not if you go by financial fundamentals. All of Space X only has around $18B in sales.

Re: Grok 4.6

#313

Earlier quoted context omitted.

Your company's owner was promoting the feature and joking about it, and called enforcement against it "fascism". CSAM generation kept up for weeks after the initial news articles, and as far as I can tell deepfake generation is still a feature. It's hard to take your AUP seriously here when you've seemingly done nothing technical to actually prevent the action.

CSAM is by definition limited to real imageries and cannot be generated. "Generative CSAM" is like "false true information". The thing about criticisms that Grok generates "CSAM" images, as well as many similar claims using that acronym, are actually more likely to be intentional mislabeling intending to refer to anime images. Advocates groups with British links love to do it, supposedly to avoid having to name state…

> CSAM is by definition limited to real imageries and cannot be generated.

Where in the definition does it imply this?

Re: Grok 4.6

#314
post #94

Earlier quoted context omitted.

Not the person you are responding to, but the fact that Grok is being used to generate a ton of CSAM and pornographic deepfakes isn't great!

(I work on Grok) This isn't allowed. CSAM / deepfakes are against our acceptable use policy.

Elon Musk literally went to court to protect the ability to make child porn with Grok:

https://www.ag.state.mn.us/Office/Communications/2026/07/31_...

Re: Grok 4.6

#315
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability.

All you need to have Fable-level AI is to announce it, and have enough fans shift from insisting that model Y is the best now, way better than model X.

Re: Grok 4.6

#316

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

It's a hack but doing things the 'proper' way is at least 1000x harder so whatever.

Is it? OpenAI released a gpt oss safeguard. You give it a policy it gives you a Rating

Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this”

https://huggingface.co/openai/gpt-oss-safeguard-120b

Re: Grok 4.6

#317
post #44
post #31

Earlier quoted context omitted.

It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now.

Does not explain timing

Yes it does, it just means all the companies come out with similar models around the same time. If what they were doing was completely novel, it would take a long time to repeat. As it is now each company releases a new model every few months, and every couple years the "leading" company changes.

Re: Grok 4.6

#318

Earlier quoted context omitted.

I believe it is because of the CEO and his recent forays into politics. The model itself is great though, especially in grok build, which is a really nice harness I find myself preferring these days.

[flagged]

Comparing Musk to Hitler is just deeply unserious.

Re: Grok 4.6

#319
post #234

Earlier quoted context omitted.

Opus 5 is terrible. I'd even say it's a step backwards from 4.8. I'm getting high error rates from it, and then it catches the error, and then it sometimes errors the error fix (!). Just today I had to switch another agent to Fable with the instruction, "Please clean up the mess that Opus 5 made, thanks" The other day, Sol called Opus 5's handoff (a skill I have that is basically a compaction, but just written to a f…

Interesting. My experience has been similar. Opus 4.8 was awesome. Opus 5 feels a little off, although I can't put my finger on exactly what it is.

Well they say Opus was trained for the subordinate role, so it doesn't excel in global view of things.

It may be a good subagent but probably not a great decision maker.

Re: Grok 4.6

#320
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

"you may find and fix vulnerabilities in local codebases only"

This seems like a bad idea, what does local mean? Anything Grok can access locally? This seems like asking for trouble.

Post reply on HN