Earlier quoted context omitted.
They will get sharply better in tasks with verifiable domains... math and coding Gradually the labs will start engineering verifiable sandboxes for wider domains like videogames This strategy will hit a plateau in about 18 months and then we're back to diminishing returns and incremental progress along other dimensions (like accelerated inference using ASICs)
You only mention math, coding and videogames. They already hire and pay people with research titles for creating and solving problems in their fields. And a lot of labs say that RL can help everywere and has plenty of way to go.
Grok 4.6
341–350 of 696 posts
Re: Grok 4.6
#342Earlier quoted context omitted.
System prompts are more like suggestions than hard constraints.
I don't understand why they don't look for large substring matches for the system prompt before returning the response. Trivial calculation compared to a system prompt instruction asking the model not to do it
Re: Grok 4.6
#343Earlier quoted context omitted.
> CSAM is by definition limited to real imageries and cannot be generated. Where in the definition does it imply this?
I thought that's just the legal definition in any sufficiently developed countries?
Re: Grok 4.6
#344Earlier quoted context omitted.
Maybe because frontier labs buy the same RL tasks from task producer companies.
Who are these task producers? Are you saying that Anthropic, et al delegate the RL part to third party companies that do it for pretty much every other AI company as well?
Re: Grok 4.6
#345Re: Grok 4.6
#346Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…
Re: Grok 4.6
#347Earlier quoted context omitted.
Can someone help me understand the deep fake controversy? That's like making photoshop illegal.
It was always possible to modify images to produce inappropriate or insensitive content, but plugging a turbocharged state of the art image generator with virtually no guardrails into every Twitter reply and then failing to address the issue long after it was obviously being used for CSAM or deepfakes of real people against their will.. well that's worse
Re: Grok 4.6
#348Earlier quoted context omitted.
Not the person you are responding to, but the fact that Grok is being used to generate a ton of CSAM and pornographic deepfakes isn't great!
It’s pretty telling that almost all of the bullet points in the system prompt that was posted for Grok have to do with preventing criminality and CSAM generation. No other provider has this same issue at that scale. The first-order-thinking reaction is “oh cool, look how they don’t want it to happen” but the second-order reaction is “why does this company have such a problem when others don’t?” It’s their own tactics…
What gives you that impression?
Re: Grok 4.6
#349Earlier quoted context omitted.
These system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output.
I think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much. "Make no mistakes"
The tooling around isolation, logging, and real time security/anonomly detection for regular LLM laptop users is very immature right now. I expect that to change soon.
The alternative is extremely locked down models which is what Anthropic seems to want to do.
Re: Grok 4.6
#350Tangental, but has anyone else noticed grok's voice mode got stupid and terse ~2 weeks ago? I've absolutely loved grok's voice mode since it came out (incredibly useful for brainstorming on walks and helping conceptualise and get the verbiage for expressing ideas) but it seems so have lost about 40 IQ points recently, and if the question is multi-part, it often answers just one part with no elaboration or explanation…