Earlier quoted context omitted.
The impression seems to be warranted: Grok 3 has directly jumpted to the top of all leaderboard categories in Chatbot Arena: https://lmarena.ai/?leaderboard In math it shares the top spot with o1 and is just a few points behind (well within errors). In creative writing it is basically ex-aequo with the latest ChatGPT 4o and in coding it's actually significantly ahead of everyone else and represents a new SOTA.
What do we do to assess the intelligence of these models after they are smarter than any human? From the kinds of questions it's answering seems like they are almost there. Do we have a way to tell if one model is smarter than another at that point?
Grok3 Launch [video]
841–850 of 1001 posts
Re: Grok3 Launch [video]
#842Earlier quoted context omitted.
I use them as a springboard for things I am really unfamiliar with. I'm self-learning electronics at the moment, and so I can ask it things like "what's a common and widely available cooperator." You will not find that answer on a search engine, I don't care how good your Google fu is. It's a weak jack of all trades: it knows a fair amount about the sum of human knowledge (which is objectively super-human), but can't…
Heads up as an EE who uses LLMs quite a bit; they cannot analyze circuits or build them. They might be able to help stitch together modules (like sensor boards plugged into microcontrollers) and definitely can write code to get things going, but they fall flat on their face hard for any kind of bare bones electronics design. Like 5% success rate and 95% totally incorrect hallucinations.
Re: Grok3 Launch [video]
#843[flagged]
“As far as a quick vibe check over ~2 hours this morning, Grok 3 + Thinking feels somewhere around the state of the art territory of OpenAI's strongest models (o1-pro, $200/month), and slightly better than DeepSeek-R1 and Gemini 2.0 Flash Thinking. Which is quite incredible considering that the team started from scratch ~1 year ago, this timescale to state of the art territory is unprecedented. Do also keep in mind the caveats - the models are stochastic and may give slightly different answers each time, and it is very early, so we'll have to wait for a lot more evaluations over a period of the next few days/weeks. The early LM arena results look quite encouraging indeed. For now, big congrats to the xAI team, they clearly have huge velocity and momentum and I am excited to add Grok 3 to my "LLM council" and hear what it thinks going forward.”
[1] Full review at: https://x.com/karpathy/status/1891720635363254772?s=46&t=91u...
Re: Grok3 Launch [video]
#844Credit to the engineers that built this, but it fills me with rage that Elon has this sort of unchecked power. How long before this starts getting deployed in safety critical applications or government decision making processes? With no oversight because Elon seems to have the power to dismiss the people responsible for investigating him. Anyone not scared by this concentration of power needs to pick up a book.
Never seen HN turned against someone so vehemently, it's as if a group of bots was set lose to criticize a certain individual.
just because you disagree with a widespread view/opinion does not mean its bots
Re: Grok3 Launch [video]
#845Grok has gotten to the top of one benchmark: https://x.com/lmarena_ai/status/1891706264800936307 It's been said before but it is great news for consumers that there's so much competition in the LLM space. If it's hard for any one player to get daylight between them & the 2nd best alternative, hopefully that means one monopolistic firm isn't going to be sucking up all the value created by these things
I think it's already clear that these are going to be commoditized and the free / open source versions will be good enough to capture enough of the value that the remaining players will not be Facebook-level monopolies on the space
Re: Grok3 Launch [video]
#846Re: Grok3 Launch [video]
#847If what they say is true, then you have to give them credit for catching up incredibly fast. And slightly pulling ahead. Not only with the models, but also products.
I have a close friend working in core research teams there. Based on our chats, the secret seems to be (1) massive compute power (2) ridiculous pay to attract top talents from established teams (3) extremelly hard work without big corp bureaucracy.
Re: Grok3 Launch [video]
#848Re: Grok3 Launch [video]
#849Earlier quoted context omitted.
It's not on par with o1, let alone o1-pro
It's on par/better/worse depending on the problem. o1 is significantly worse, for example, in Rust programming than Claude 3.5; at least for me.
Re: Grok3 Launch [video]
#850Have you thought of a future where LLM will be fined tune to target advertisment to you? I mean look at search: first iterations of search were pretty simple in term of ads. Then personalized ads came. I wouldn't help but envision the distopia where the LLM will insert personalized ads based on what you are asking for help.
It's way worse than that. First, We interact with LLMs through private conversation and we are used to have private conversation with human we trust. Some of that trust will be transfered to LLMs. Second, LLMs have a vastly bigger "mental" power to build a long term mental model of us, while we interact with them. Which mean they can chose with extreme precision their words to trigger an emotion, a certain reaction.…