[flagged]
Grok3 Launch [video]
871–880 of 1001 posts
Re: Grok3 Launch [video]
#872[flagged]
Re: Grok3 Launch [video]
#873[flagged]
Re: Grok3 Launch [video]
#874A very impressive debut. No doubt they benefited from all the research and discoveries that have preceded it. Maybe the best outcome of a competitive Grok is breaking the mindshare stranglehold that ChatGPT has on the public at large and with HN. There are many good frontier models that are all very close in capabilities.
Re: Grok3 Launch [video]
#875Earlier quoted context omitted.
The impression seems to be warranted: Grok 3 has directly jumpted to the top of all leaderboard categories in Chatbot Arena: https://lmarena.ai/?leaderboard In math it shares the top spot with o1 and is just a few points behind (well within errors). In creative writing it is basically ex-aequo with the latest ChatGPT 4o and in coding it's actually significantly ahead of everyone else and represents a new SOTA.
What do we do to assess the intelligence of these models after they are smarter than any human? From the kinds of questions it's answering seems like they are almost there. Do we have a way to tell if one model is smarter than another at that point?
Ask them to design a ranking mechanism for you. They are superhuman, after all.
(I really don't think we're going to have to worry about this).
Re: Grok3 Launch [video]
#876Earlier quoted context omitted.
It is a debut of their thinking mode iirc. Unfortunately LLMs are shifting compute time to test time instead of train time. I don't really like this and frankly it shows a stalling of the architectures, data sets, etc...
Another take is that the base models are now good enough that spending more money for more intelligence is viable at test time. A threshold has been crossed.
Naively, I feel to be useful, the goal of LLMs should be to more power efficient. So that eventually all devices can be smarter.
Power efficiency can be gained through less time-time, or more "intelligence" or some combination of the two. I'm not convinced these SOTA models are doing much more than increasing test-time.
Re: Grok3 Launch [video]
#877Earlier quoted context omitted.
The full-sized DeepSeek-R1 is on par with o1. o1-pro is "o1 on steroids" and was the first selling point of the $200/month Pro subscription but they later also added "Deep Research" and Operator to the Pro subscription.
Every year seems like we get worse at naming things in non confusing ways. I am waiting for the o1-pro-max now, pro max ultra and pro max ultra plus.
Re: Grok3 Launch [video]
#878[flagged]
People who buy into Grok are willingly submitting themselves to the far-right propaganda machine. I’m sure it’s nice and tidied up for release, but there is zero chance that Musk will not use this tool to push his ideological agenda given its reach and impact.
Re: Grok3 Launch [video]
#879Re: Grok3 Launch [video]
#880[flagged]
Here is just one headline from today, The Elon Musk-led Department of Government Efficiency (DOGE) on Monday revealed its finding that $4.7 trillion in disbursements by the US Treasury are "almost impossible" to trace, thanks to a rampant disregard for the basic accounting practice of using of tracking codes when dishing out money.