Ok speed (202.7 tok/s) and value (1.25 -> 2.50) look great, with pretty decent intelligence.
[0]: https://aibenchy.com/compare/x-ai-grok-4-20-medium/x-ai-grok...
451–460 of 608 posts
Ok speed (202.7 tok/s) and value (1.25 -> 2.50) look great, with pretty decent intelligence.
[0]: https://aibenchy.com/compare/x-ai-grok-4-20-medium/x-ai-grok...
Do people really use Grok for anything outside of Twitter memes or understanding tweets? I'm asking out of genuine curiosity.
https://arstechnica.com/tech-policy/2026/03/elon-musks-xai-s...
Earlier quoted context omitted.
I know it's just an evaluation, but seeing an informal message and a prompt to ask to rewrite this informal message to the tone of an "informal message" when the original one sounds just fine, just makes me sad... Not because of this evaluation, but because it reminds me that this is how some people use LLMs, basically asking it to remove your own voice from texts that are generally fine already.
My sister in law is a pharmacist and the heaviest non-dev ChatGPT user I know and her main use case is writing professionally polite messages to doctors on how the drugs they prescribed to a patient would have killed them had she not caught a particular interaction or common side effect. There's a lot of "tone" in it as she's not trying to anger these folks, but also it's quite serious, but also there's just everythi…
Earlier quoted context omitted.
Credit where it's due, Grok is currently the only model that has near-realtime updates from/access to a waterhose of data, and is casually used by regular people all the time. I don't think there's a single thread on Xitter whete people don't delegate some question to grok. (There's a separate conversation of failure modes, and whether it's a good thing, and how much control Elon had when he doesn't like Grok's "woke…
All the major tools can websearch guy
Earlier quoted context omitted.
not really. there are easy heuristics to filter out bots with good confidence. FWIW i don't see any bots posting anything in my feed
congratulations, you have solved anti-scam. go make your billion since its easy.
you think its hard?
Earlier quoted context omitted.
This puts Sonnet 4.6 above Opus 4.6 in the coding index.. kinda hard to trust those numbers. (Also it puts Opus 4.7 universally above Opus 4.6, and I may be wrong but this doesn't seem to match the experience of most/many/some people. I think it's widely recognized that Anthropic is severely lacking compute and Opus 4.7 is a costs saving measure)
What I’ve usually seen is 4.7 -> 4.5 -> 4.6 in terms of quality. Though 4.7 seems to hallucinate more than before.
Earlier quoted context omitted.
It's being biased on purpose. Musk has intervened multiple times when he believed Grok's responses were too "woke" or "leftist". https://www.nytimes.com/2025/09/02/technology/elon-musk-grok... In response to Grok saying that the "woke mind virus is often exaggerated" the prompt was tweaked so that Grok now says "The woke mind virus 'poses significant risks'" If you truly believed in what your comment states then you…
The new response works for me, because in my mind I’ve always defined “woke mind virus” as a a mental virus which causes people to become absolutely pathologically obsessed with fighting an imaginary enemy they call “wokeness”. It’s the only definition which makes sense. “Woke” itself was never that viral.
People obsessed with fighting whatever they perceive as "woke" which remains ill-defined on purpose so they never have to actually formulate a rational take down beyond their emotional response
Earlier quoted context omitted.
Do you drive BMW or VW car? Boy do I have news for you!
Go on...make your case
Do people really use Grok for anything outside of Twitter memes or understanding tweets? I'm asking out of genuine curiosity.
[0] sometimes you need to lightly jailbreak it, or rerun the prompt, the non-deterministic nature means sometimes you will get a refusal