Grok 4.1
31–40 of 135 posts
Re: Grok 4.1
#32This model has effectively no safety filters (even fewer than Grok 4 in my testing), which I've confirmed via this web release: https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib... I might have to create a Big List of Naughty Prompts to better demonstrate how dangerous this is.
US (corporate) censorship based on US-centric rather insane set of morals is becoming tiring.
Re: Grok 4.1
#33https://tools.simonwillison.net/svg-render#%3Csvg%20width%3D...
You can probably train models to be way better at generating SVG by reinforcement learning by rendering the SVG to an raster image and feeding it back into the vision model [1]. Same with, say, generating HTML/CSS webpages. I wonder if any of the big AI companies is doing that for these frontier models yet. [1] https://arxiv.org/abs/2505.20793
Re: Grok 4.1
#34This is generally a challenging prompt for LLMs - it requires knowledge of the story, ideally the LLM would have seen the Roseanne Barr video, not just read about it in the New Yorker. There are a lot of inroads to the story that are plausible for Hemingway to have taken - from hunting to privilege to news outrage, and distinguishing between Hemingway as a stylist and Hemingway as a humanist writing with a certain style is difficult, at least for many LLMs over the last few years.
Grok 4.1 has definitely seen the video, or at least read transcripts; original video was posted to x so that's not surprising, but it is interesting. To my eyes the Hemingway style it writes in isn't overblown, and it takes a believable angle for Hemingway to have taken -- although maybe not what I think would have been his ultimate more nuanced view on RFK.
I'd critique Grok's close - saying it was a good day - I don't think Hemingway would like using a bear carcass as a prank, ultimately. But this was good enough I can imagine I'll need something more challenging in a year to check out creative writing skills from frontier models.
https://grok.com/share/bGVnYWN5LWNvcHk_92bf5248-18e1-4f8a-88...
Re: Grok 4.1
#35It seems Grok 4.1 uses more emojis than 4.
Also GPT5.1 thinking is now using emojis, even in math reasoning. 5 didn't do that.
Re: Grok 4.1
#36Not a big fan of emojis becoming the norm in LLM output. It seems Grok 4.1 uses more emojis than 4. Also GPT5.1 thinking is now using emojis, even in math reasoning. 5 didn't do that.
…but I’m still infuriated when I read a passage full of them.
Re: Grok 4.1
#37Not a big fan of emojis becoming the norm in LLM output. It seems Grok 4.1 uses more emojis than 4. Also GPT5.1 thinking is now using emojis, even in math reasoning. 5 didn't do that.
Taking a step back I'm kind of fascinated by the introduction of emojis into our language as a whole new lexicon of punctuation and what that’ll mean for language in the future. …but I’m still infuriated when I read a passage full of them.
Re: Grok 4.1
#38https://tools.simonwillison.net/svg-render#%3Csvg%20width%3D...
Huh, it decided to drop in a seal and bike emoji? What happens if you ask it if a seahorse emoji exists?
https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58...
Although it does eventually come to the right conclusion... sort of.
Re: Grok 4.1
#39No mention of coding benchmarks. I guess they've given up on competing with Claude and GPT-5 there. (and from my initial testing of grok 4.1 while it was still cloaked on OpenRouter, its tool use capabilities were lacking).
Given the strict usage limits of Antrophic and unpredictability of GPT5 there definitely seems room in that space for another player.
Re: Grok 4.1
#40Not a big fan of emojis becoming the norm in LLM output. It seems Grok 4.1 uses more emojis than 4. Also GPT5.1 thinking is now using emojis, even in math reasoning. 5 didn't do that.
> Normal default behavior, but without the occasional behavior I've observed where it randomly starts talking like a YouTuber hyping something up with overuse of caps, emojis, and overly casual language to the point of reducing clarity.