https://artificialanalysis.ai/models/grok-4-3
It says #1 for speed but then in the chart it's #2. Also says #10 for intelligence but then it's #7 in the chart.
11–20 of 608 posts
https://artificialanalysis.ai/models/grok-4-3
It says #1 for speed but then in the chart it's #2. Also says #10 for intelligence but then it's #7 in the chart.
Ok speed (202.7 tok/s) and value (1.25 -> 2.50) look great, with pretty decent intelligence.
Pelican riding a bike here: https://gist.github.com/SerJaimeLannister/f6de26bd0d0817e056... (ran this on arena.ai direct chat and also tried to write this gist inspired by how simon writes his gists about pelicans) Edit: just realized that I made pelican riding a bike instead of bicycle, which now makes sense as to why it hardened the bicycle to look tankier, going to compare this with pelican riding a bicycle if any…
You should probably come up with variations, like a beaver riding a scooter or something, just to see what's what :)
Expensive miscalculation.
Just wish they would finally put some work into their apps, it's the only thing keeping me from actually subscribing to SuperGrok:
- No MCP / connected apps support. It's been teased but here we are, still not available. I can't connect Grok to anything, so I can't use it for serious work
- Projects are still not available in the app so as soon as you move something into a project, it's gone from all the native apps
- No way to add artifacts (like generated markdown docs) directly to a project, we have to export to PDF/markdown and re-import. And there isn't even a way to export artifacts. This makes serious project work hard because we can't dynamically evolve projects with new information
- No memory, no ability to look up other chats, each chat is completely new
- No voice mode in projects at all
If someone from xAI is reading this, please consider adding some of these.
Ok speed (202.7 tok/s) and value (1.25 -> 2.50) look great, with pretty decent intelligence.
The problem with speed is that they usually are very fast for first few weeks and then suddenly much slower. They did such trick when they advertised Grok 4 fast ( dropped from 200 tps to 60tps)
When looking at the benchmarks, this model seems to be really close to Kimi K2.6 in terms of intelligence and pricing, hitting that sweet spot. It does also have a higher AA-Omniscience index, which is something kimi and other open models lack in. Curious to see how pleasant it is to use.