Live data from Hacker News

Grok 4.6

x.ai

301–310 of 696 posts

Re: Grok 4.6

#301
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

> 2) Distillation - also implausible for the reason above.

DeepSeek V4 Flash 0731 is a distilled version of Fable into the original V4 Flash (announced before Fable), to the point that it also says load bearing and what not.

Re: Grok 4.6

#303
post #95
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually. What are suspicious of? If the timing is similar maybe just everyone already are of similar capabilities and got there at a similar time? > Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? It means Anthropic had no…

Well, Opus 5 and Fable are the only models I don’t constantly swear at and call stupid, which seems like a pretty good moat to me.

My guess is all the commenters (you are the 4th person I’ve seen say this) saying ‘Anthropic has no moat’ haven’t actually used Fable or even Opus 5 yet. Sol is laughable by comparison, and Grok… lol.

Re: Grok 4.6

#304
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It's just model size and heavy RL, sometimes they overfit on specific tasks. RL can get you very far, prior models did not have such a focus on RL for agentic setups.

Look at deepseek, they improved it just by doing a lot of RL and you can see it from how it behaves. You provide very little information about a task, but since they are trained on similar tasks, they come up with a lot of assumptions and details on their own, because they were trained with such an info during RL.

Re: Grok 4.6

#305
post #290
post #44

Earlier quoted context omitted.

Does not explain timing

Maybe because frontier labs buy the same RL tasks from task producer companies.

Who are these task producers? Are you saying that Anthropic, et al delegate the RL part to third party companies that do it for pretty much every other AI company as well?

Re: Grok 4.6

#306
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It was said at the time that xAI acquiring Cursor was very smart because it would give them access to years of agent coding traces from millions of users.

$60B in SpaceX stock for Cursor was a bargain

Data + compute + being competent and smart enough to ship.

fwiw I don't think these are yet Fable level - the difference tends to get discovered in the long tail of tasks - but they're close enough, they're cheap, and the length of the frontier exclusive window is narrowing

Re: Grok 4.6

#307

Earlier quoted context omitted.

I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…

The US govt trusts SpaceXAI for defense and high security missions. The idea they are lying about contracted AI services is absurd. They're also a public company which beings even more oversight than openai / anthropic.

Their closeness to the current US government is a cause for concern, it doesn't alleviate it.

Re: Grok 4.6

#308
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

It's a hack but doing things the 'proper' way is at least 1000x harder so whatever.

Re: Grok 4.6

#309

Earlier quoted context omitted.

I'm sure the SF AI scene leaks like a sieve, and companies have a pretty good idea what each other is working on.

Okay so everyone is blaming diffusion or spying or whatever but we all use all of the models on our various projects in aggregate and they get to all read the code each other is generating. I do this with research tasks and local random stuff too. So why do people have this idea in their heads that it's all some sorta secret sauce they are taking from each other?

I didn't mean that - I meant that when, for example, Anthropic started, then later finished their Mythos/Fable pre-training run that people at OpenAI and elsewhere would have heard about it, probably knew some details such as the size of the model etc - people from these companies go out and socialize with each other, attend parties, share houses ...

So, it's not coincidence when they respond to each others models with something roughly equivalent - because they know what each other are working on.

Re: Grok 4.6

#310

Earlier quoted context omitted.

i beg to differ, in an ideal world a system possibly is a binding law and high end models are starting to be really aligned to the exact system prompt. The instructions must be simple to follow, if you start doing complex rules it'll call apart, but I'll usually follow the stringer interpretation.

"I beg to differ, it is my opinion that reality should be different to what you have observed"

in reality even the mention of a prohibition is enough to make the model reject that no matter what
Post reply on HN