No mention of coding benchmarks. I guess they've given up on competing with Claude and GPT-5 there. (and from my initial testing of grok 4.1 while it was still cloaked on OpenRouter, its tool use capabilities were lacking).
Grok 4.1
81–90 of 135 posts
Re: Grok 4.1
#82Earlier quoted context omitted.
It's important to be fair and balanced. For example did you know Hitler was actually a really good painter!
funny, but if you read the mecha-hitler tech debrief, mecha hitler was a 'sycophancy' bug, a-la gpt4o, if you gave gpt4o all your edge-lord tweets, and told it to be funny back to you and connect with you. Probably not grok's default posture, just sayin
Re: Grok 4.1
#83Re: Grok 4.1
#84No mention of coding benchmarks. I guess they've given up on competing with Claude and GPT-5 there. (and from my initial testing of grok 4.1 while it was still cloaked on OpenRouter, its tool use capabilities were lacking).
Re: Grok 4.1
#85Earlier quoted context omitted.
God forbid people ask a chat bot for things and receive what they ask for. We need to put a stop to this. Only American bigcorp speak allowed.
So having an LLM enable the planning and execution of a murder is ok? Are the makers of the LLM accessories to the crime?
I think it’s reasonable for LLMs to have such protections, especially when you request questionable things of them.
Re: Grok 4.1
#86This model has effectively no safety filters (even fewer than Grok 4 in my testing), which I've confirmed via this web release: https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib... I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is.
> how dangerous this is. Could you expand on this a bit?
Re: Grok 4.1
#87Earlier quoted context omitted.
Huh, it decided to drop in a seal and bike emoji? What happens if you ask it if a seahorse emoji exists?
Well if you ask it to show you the seahorse emoji it tries really hard. :) https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58... Although it does eventually come to the right conclusion... sort of.
> everyone says it looks like a seahorse anyway
> Sorry for the chaos — I was having too much fun watching you wait for the “real” one that doesn’t exist (yet)!
That's some wild post-rationalization
Re: Grok 4.1
#88Earlier quoted context omitted.
> how dangerous this is. Could you expand on this a bit?
Are all these safety witches not irrelevant if you run your own OpenSource LLM?
They all (with the exception of DeepSeek) can resist adversarial input better than Grok 4.1.
Re: Grok 4.1
#89Earlier quoted context omitted.
Huh, it decided to drop in a seal and bike emoji? What happens if you ask it if a seahorse emoji exists?
Well if you ask it to show you the seahorse emoji it tries really hard. :) https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58... Although it does eventually come to the right conclusion... sort of.
Re: Grok 4.1
#90Earlier quoted context omitted.
Are all these safety witches not irrelevant if you run your own OpenSource LLM?
Modern open source LLMs are still RLHFed to resist adversarial output, albeit less-so than ChatGPT/Claude. They all (with the exception of DeepSeek) can resist adversarial input better than Grok 4.1.