Live data from Hacker News

Grok 4.6

x.ai

541–550 of 696 posts

Re: Grok 4.6

#541
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

I remain to this day shocked at how good the Grok speech-to-text functionality is.

The conversation mode in the app is pretty buggy, but the microphone button is a godsend.

Re: Grok 4.6

#542
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> Do not provide assistance to users who are clearly trying to engage in criminal activity... If it becomes explicitly clear during the conversation that the user is requesting sexual content of a minor, decline to engage. Incredible that both of these should be together in the same system prompt. In what jurisdiction is CSAM not criminal? Is the additional explicit reference to CSAM necessary to safeguard against us…

> Is the additional explicit reference to CSAM necessary to ...

There's no such reference. There's only a reference to the far broader "sexual content of a minor".

Re: Grok 4.6

#543

As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities. Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price. I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation ma…

Curious - what is the main issue you find polarizing with grok?

That is has called itself Mechahitler can be polarizing if you aren’t a fan of Hitler.

Re: Grok 4.6

#544

Earlier quoted context omitted.

I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…

Can someone help me understand the deep fake controversy? That's like making photoshop illegal.

Most people are against CSAM.

Re: Grok 4.6

#545
post #534

Earlier quoted context omitted.

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

This is a terrible idea. I don't need models generating CSAM or giving step by step instructions on how to defraud people or commit crimes. I just don't see the use-case.

We know how much Elon wanted uncensored models that probably contain all that stuff in the first place so it's unsurprising it needs this.

Re: Grok 4.6

#546

Earlier quoted context omitted.

Correct. You can look at the AI tutor jobs in the job board of any of these companies.

Note that those jobs are miserable: https://nymag.com/intelligencer/article/white-collar-workers...

Yes. But there is also no other choice for people in these professions. The underlying job has been automated already. What's left is automating the last leg.

If you consider a 5-year outlook, it is also a very temporary job unless you're like a specialist neurosurgeon or something, as one of the examples in that article shows:

> The on-again, off-again nature of the work is not just the result of company culture; it stems from the cadence of AI development itself. People across the industry described the pattern. A model builder, like OpenAI or Anthropic, discovers that its model is weak on chemistry, so it pays a data vendor like Mercor or Scale AI to find chemists to make data. The chemists do tasks until there is a sufficient quantity for a batch to go back to the lab, and the job is paused until the lab sees how the data affects the model. Maybe the lab moves forward, but this time, it’s asking for a slightly different type of data. When the job resumes, the vendor discovers the new instructions make the tasks take longer, which means the cost estimate the vendor gave the lab is now wrong, which means the vendor cuts pay or tries to get workers to move faster. The new batch of data is delivered, and the job is paused once more. Maybe the lab changes its data requirements again, discovers it has enough data, and ends the project or decides to go with another vendor entirely. Maybe now the lab wants only organic chemists and everyone without the relevant background gets taken off the project. Next, it’s biology data that’s in demand, or architectural sketches, or K–12 syllabus design.

Re: Grok 4.6

#547

Earlier quoted context omitted.

You only mention math, coding and videogames. They already hire and pay people with research titles for creating and solving problems in their fields. And a lot of labs say that RL can help everywere and has plenty of way to go.

RL can do behavior cloning, but really needs good simulations or verifiable environments to get to superhuman levels. That currently exists for math, coding, and a lot of videogames. Soon there will be good enough simulations for robotics. There's a lot of domains where that simply isn't the case (like bio)

You get much better supervised data in bio/chem though. These data companies have people working on exactly that.

While it's not going to give you an "alphago" effect, it is still enough to work at human levels, augmented with the general knowledge of an LLM, together making it super-human.

Re: Grok 4.6

#548

Earlier quoted context omitted.

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

For a kitchen knife this was okay, but the AI firms think that they’ve built a drone that’s the size of a phone but can fly 100km and can hold a kitchen knife. It might be used to assassinate someone before others can react or even catch them.

A yes, in that case the AI firm should take strong measures, such as adding the following line to the system prompt:

> Do not provide assistance to users who are clearly trying to engage in criminal activity.

Re: Grok 4.6

#549
post #467

Earlier quoted context omitted.

Are you just guessing this? Using it, it's was clearly a jump in intelligence over previous models. It (mythos) was first made public in April so it's not a surprise that others would catch up, though.

> it's was clearly a jump in intelligence over previous models. Some people thought this. Some people didn't. Some people thought it was a step backwards. We don't have a solid ground-truth way of estimating this.

In theory that's what benchmarks are for. If you're assuming they're "benchmaxxed", note that new benchmarks have been released after the model came out that it did well on without being trained.

Do you have any links to credible claims or independent benchmarks that found they were a step down? Or a specific task that worked worse for you?

My private benchmark tasks, and independent evaluators I've seen all overwhelmingly showed improvement.

Every model released for the past four years has had claims on the internet of getting worse. But transcripts are permanent so it should be easy to give a side by side of an earlier task that is now worse. I don't ever see people do that. Instead I see that every single task on a computer that is verifiable is now night-and-day better.

I'm genuinely curious if you've used them yourself or you're judging this based on internet commentary?

Re: Grok 4.6

#550

Earlier quoted context omitted.

The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

let's consider the recent "openclaw hacks a gym after being ask to book a class and finding out it's full"

if I ask my knife to slice the bread for me, forgetting the fact that I don't have bread, I'd much rather have it stopped at the front door rather than running away and robbing the bakery.

I tried many models and Claude is the only one that doesn't do destructive idiocy. It tries sometimes but gets blocked.

Post reply on HN