Live data from Hacker News

Grok 4.1

x.ai

81–90 of 135 posts

Re: Grok 4.1

#81

No mention of coding benchmarks. I guess they've given up on competing with Claude and GPT-5 there. (and from my initial testing of grok 4.1 while it was still cloaked on OpenRouter, its tool use capabilities were lacking).

I've often used Grok Heavy to get me past a problem when Claude gets stuck. Not always, but it usually can figure it out.

Re: Grok 4.1

#82
post #26

Earlier quoted context omitted.

It's important to be fair and balanced. For example did you know Hitler was actually a really good painter!

funny, but if you read the mecha-hitler tech debrief, mecha hitler was a 'sycophancy' bug, a-la gpt4o, if you gave gpt4o all your edge-lord tweets, and told it to be funny back to you and connect with you. Probably not grok's default posture, just sayin

but but hivemind

Re: Grok 4.1

#84

No mention of coding benchmarks. I guess they've given up on competing with Claude and GPT-5 there. (and from my initial testing of grok 4.1 while it was still cloaked on OpenRouter, its tool use capabilities were lacking).

They've got Grok Code Fast. Maybe they want to split than out from the general purpose model.

Re: Grok 4.1

#85
post #73
post #29

Earlier quoted context omitted.

God forbid people ask a chat bot for things and receive what they ask for. We need to put a stop to this. Only American bigcorp speak allowed.

So having an LLM enable the planning and execution of a murder is ok? Are the makers of the LLM accessories to the crime?

As you’re on this platform, you’re a beneficiary of Section 230 protections.

I think it’s reasonable for LLMs to have such protections, especially when you request questionable things of them.

Re: Grok 4.1

#86
post #48

This model has effectively no safety filters (even fewer than Grok 4 in my testing), which I've confirmed via this web release: https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib... I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is.

> how dangerous this is. Could you expand on this a bit?

Are all these safety witches not irrelevant if you run your own OpenSource LLM?

Re: Grok 4.1

#87
post #38

Earlier quoted context omitted.

Huh, it decided to drop in a seal and bike emoji? What happens if you ask it if a seahorse emoji exists?

Well if you ask it to show you the seahorse emoji it tries really hard. :) https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58... Although it does eventually come to the right conclusion... sort of.

> I swear this one looks like a tiny seahorse when you squint

> everyone says it looks like a seahorse anyway

> Sorry for the chaos — I was having too much fun watching you wait for the “real” one that doesn’t exist (yet)!

That's some wild post-rationalization

Re: Grok 4.1

#88
post #48

Earlier quoted context omitted.

> how dangerous this is. Could you expand on this a bit?

Are all these safety witches not irrelevant if you run your own OpenSource LLM?

Modern open source LLMs are still RLHFed to resist adversarial output, albeit less-so than ChatGPT/Claude.

They all (with the exception of DeepSeek) can resist adversarial input better than Grok 4.1.

Re: Grok 4.1

#89
post #38

Earlier quoted context omitted.

Huh, it decided to drop in a seal and bike emoji? What happens if you ask it if a seahorse emoji exists?

Well if you ask it to show you the seahorse emoji it tries really hard. :) https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58... Although it does eventually come to the right conclusion... sort of.

Now we get to guess if it's broken in the same way as gpt, or did it pick up that pattern from all the cases of people posting it on the internet. (In the second case, that's not a good look for their data cleanup process)

Re: Grok 4.1

#90

Earlier quoted context omitted.

Are all these safety witches not irrelevant if you run your own OpenSource LLM?

Modern open source LLMs are still RLHFed to resist adversarial output, albeit less-so than ChatGPT/Claude. They all (with the exception of DeepSeek) can resist adversarial input better than Grok 4.1.

Is this not easy to take out/deactivate?
Post reply on HN