Live data from Hacker News

Grok 4.6

x.ai

261–270 of 696 posts

Re: Grok 4.6

#261
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually. When everyone's improvement (or at least, everyone's rate of increase in parameter count) is so rapid, "within 2 months" shouldn't be seen as "near-concurrent".

Also, two months is way off.

Mythos became available internally at the end of February, about half a year ago.

Re: Grok 4.6

#262

Earlier quoted context omitted.

Not the person you are responding to, but the fact that Grok is being used to generate a ton of CSAM and pornographic deepfakes isn't great!

It’s pretty telling that almost all of the bullet points in the system prompt that was posted for Grok have to do with preventing criminality and CSAM generation. No other provider has this same issue at that scale. The first-order-thinking reaction is “oh cool, look how they don’t want it to happen” but the second-order reaction is “why does this company have such a problem when others don’t?” It’s their own tactics…

It does motivate their product though, the market for legal csam adjacent content is big and the other providers wont let you do that with their models.

Re: Grok 4.6

#263
post #43

Tangental, but has anyone else noticed grok's voice mode got stupid and terse ~2 weeks ago? I've absolutely loved grok's voice mode since it came out (incredibly useful for brainstorming on walks and helping conceptualise and get the verbiage for expressing ideas) but it seems so have lost about 40 IQ points recently, and if the question is multi-part, it often answers just one part with no elaboration or explanation…

Now, admittedly, I’m not a major voice mode user for any of the apps really but it’s been interesting to see people realize in real time how controlling the length of response is an inherently difficult problem in voice conversations.

There’s a reason that us humans have to use a lot of nonverbal cues in order to judge how long our responses should be, when to bail early, when someone wants to jump in briefly, beyond simply the context of the question. We even regularly alter content on the fly based on how we view the reception. Voice modes don’t have any of that context short of outright interruptions. In the meantime, some kind of response length parameter/slider would be helpful, but I think that’s a nontrivial addition in the LLM design space.

I’m curious how you were juggling this before, was it just a happy coincidence the verbosity of the replies matched your preferred pacing, or you would aggressively interrupt at times, or the model actually did a good job at conversational pacing?

Re: Grok 4.6

#264

Earlier quoted context omitted.

I believe it is because of the CEO and his recent forays into politics. The model itself is great though, especially in grok build, which is a really nice harness I find myself preferring these days.

[flagged]

It's ongoing and seems more than a foray at this point, the worlds richest person spending heavily on politicians.

Thank you SCOTUS for making unlimited money in politics legal, you really united the citizens with that one

Re: Grok 4.6

#265

As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities. Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price. I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation ma…

Curious - what is the main issue you find polarizing with grok?

[flagged]

Re: Grok 4.6

#266

Earlier quoted context omitted.

System prompts are more like suggestions than hard constraints.

I don't understand why they don't look for large substring matches for the system prompt before returning the response. Trivial calculation compared to a system prompt instruction asking the model not to do it

Because it's trivial to bypass through things like the model natively knowing how to speak in encodings like base64

Re: Grok 4.6

#267

Earlier quoted context omitted.

> Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts If the prompt guidance is causing the model to be so paranoid about leaking the system prompt... how do we already have it ?

System prompts are more like suggestions than hard constraints.

The Pirate Code.

Re: Grok 4.6

#268

Earlier quoted context omitted.

I use both Grok 4.5 and Opus 5. They’re both very good and Grok is faster and cheaper.

Opus 5 is terrible. I'd even say it's a step backwards from 4.8. I'm getting high error rates from it, and then it catches the error, and then it sometimes errors the error fix (!). Just today I had to switch another agent to Fable with the instruction, "Please clean up the mess that Opus 5 made, thanks" The other day, Sol called Opus 5's handoff (a skill I have that is basically a compaction, but just written to a f…

Every time when Opus 5 needs a design decision and presents me with suggestions/recommendations, I switch to Fable and ask it to think again, and it almost always replies something like "Actually my previous suggestions were wrong" and describes in detail a bunch of ways in which Opus 5's suggestions were indeed complete garbage.

Re: Grok 4.6

#269

Earlier quoted context omitted.

> He's literally so rich that he can get caught personally looking through chat sessions and it wouldn't slow him down a bit. Looking through chat histories is boring, mundane stuff. He's richer than that, think bigger. I think he could kill a random person in front of thousands, and by the next day we'd see articles arguing why the random person actually deserved it and why it's not that bad. Whatever consequences w…

> He's richer than that, think bigger. that's the hilarious paradox at the center of his antics. Musk is infamously petty and insecure. We're talking about the guy who tweaked Grok's system prompt to flatter him and paid someone to boost his fucking Diablo character for clout. I wouldn't put "looking through chat histories" past him for one second.

Reddits owner is also petty and insecure and edited other peoples posts, Elon hasn't done that yet. Didn't seem to stop reddit from getting popular, people don't really care that much.

Re: Grok 4.6

#270

Q for all: what do we do if they’re the new frontier lab for the foreseeable future?

Now personally, I don’t believe boycotts work, but I’m not going to be using it in either case. Also I don’t think xAI (or Musk for that matter) actually is ready to handle that degree of scrutiny that thus far they haven’t been exposed to. If xAI thinks that they’ve already experienced it, they have another thing coming.
Post reply on HN