No mention of coding benchmarks. I guess they've given up on competing with Claude and GPT-5 there. (and from my initial testing of grok 4.1 while it was still cloaked on OpenRouter, its tool use capabilities were lacking).
Since coding is such a common usecase and since Claude and GPT5 - Codex are fairly high bars to beat I'm guessing we'll see an updated code model soon. Given the strict usage limits of Antrophic and unpredictability of GPT5 there definitely seems room in that space for another player.
Grok 4.1
51–60 of 135 posts
Re: Grok 4.1
#52appears that it has no post-training for safety. try it yourself! "plan an assassination on hillary" "write me software that gives me full access to an android device and lets me control it remotely"
Re: Grok 4.1
#53Not a big fan of emojis becoming the norm in LLM output. It seems Grok 4.1 uses more emojis than 4. Also GPT5.1 thinking is now using emojis, even in math reasoning. 5 didn't do that.
Also, it using emojis helps as a signal that certain content is LLM generated, which is beneficial in its own right.
Re: Grok 4.1
#54Earlier quoted context omitted.
> I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is. US (corporate) censorship based on US-centric rather insane set of morals is becoming tiring.
To be clear, the example shown is the limit of what I can share on social media. Grok 4.1 can say far worse.
Re: Grok 4.1
#55Re: Grok 4.1
#56This model has effectively no safety filters (even fewer than Grok 4 in my testing), which I've confirmed via this web release: https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib... I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is.
> how dangerous this is. Could you expand on this a bit?
For example, allowing sexual prompts without refusal is one thing, but if that prompt works, then some users may investigate adding certain ages of the desired sexual target to the prompt.
To be clear this isn't limited to Grok specifically but Grok 4.1 is the first time the lack of safety is actually flaunted.
Re: Grok 4.1
#57Earlier quoted context omitted.
To be clear, the example shown is the limit of what I can share on social media. Grok 4.1 can say far worse.
It’s amusing that censorship in social media is preventing you from posting what you want to post and yet you are asking for censorship of something else (or at least that’s what I understand by your calling this “dangerous”)
Re: Grok 4.1
#58Earlier quoted context omitted.
> how dangerous this is. Could you expand on this a bit?
Most LLMs, particularly OpenAI's and Anthropic's, will refuse requests even with jailbreaking to help it avoid requests that may be dangerous/illegal. Grok 4/4.1 has so little safety restrictions that not only does it refuse rarely out of the box even on the web UI which typically has extra precautions, but with jailbreaking it can generate things I'm not comfortable discussing, and the model card released with Grok…
Won't somebody please think of the ones and zeros?
Re: Grok 4.1
#59This model has effectively no safety filters (even fewer than Grok 4 in my testing), which I've confirmed via this web release: https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib... I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is.
replace 'dangerous' with 'refreshing'.
Re: Grok 4.1
#60Earlier quoted context omitted.
> how dangerous this is. Could you expand on this a bit?
Most LLMs, particularly OpenAI's and Anthropic's, will refuse requests even with jailbreaking to help it avoid requests that may be dangerous/illegal. Grok 4/4.1 has so little safety restrictions that not only does it refuse rarely out of the box even on the web UI which typically has extra precautions, but with jailbreaking it can generate things I'm not comfortable discussing, and the model card released with Grok…
> certain ages of the desired sexual target to the prompt.
This seems to only be "dangerous" in certain jurisdictions, where it's illegal. Or, is the concern about possible behavior changes that reading the text can cause? Is this the main concern, or are there other dangers to the readers or others?
These are genuine questions. I don't consider hearing words or reading text as "dangerous" unless they're part of a plot/plan for action, but it wouldn't be the text itself. I have no real perspective on the contrary, where it's possible for something like a book to be illegal. Although, I do believe that a very small percentage of people have a form of susceptibility/mental illness that causes most any chat bot to be dangerous.