Live data from Hacker News

Grok3 Launch [video]

x.com

831–840 of 1001 posts

Re: Grok3 Launch [video]

#831

Karpathy gave his initial impression: https://x.com/karpathy/status/1891720635363254772 The pull quote is: The impression overall I got here is that this is somewhere around (OpenAI) o1-pro capability

The impression seems to be warranted: Grok 3 has directly jumpted to the top of all leaderboard categories in Chatbot Arena: https://lmarena.ai/?leaderboard In math it shares the top spot with o1 and is just a few points behind (well within errors). In creative writing it is basically ex-aequo with the latest ChatGPT 4o and in coding it's actually significantly ahead of everyone else and represents a new SOTA.

What do we do to assess the intelligence of these models after they are smarter than any human? From the kinds of questions it's answering seems like they are almost there.

Do we have a way to tell if one model is smarter than another at that point?

Re: Grok3 Launch [video]

#832
post #447

Earlier quoted context omitted.

the IP rights holders have yet to bare their teeth. I don't think the outcome you suggest is clear at all, in fact I think if anything entirely the opposite is the most probable outcome. I've lost count of the number of technology epochs that at the time were either silently or explicitly dependent on ignoring the warez aspects while being blinded by the possibilities, Internet video, music and film all went through…

>I quite like the idea of a future where the AI job holocaust largely never happened because license costs ate up most of the innovation benefit. Not quite realistic. You are talking about very huge benefits, in favor of which licenses will be abandoned. And who don't abandoned them... I mean you can look at the Amish settlements.

I'd put solid money on Warner earning a few cents every time an AI girlfriend somewhere sings happy birthday within 10 years

Re: Grok3 Launch [video]

#833
post #10

Grok has gotten to the top of one benchmark: https://x.com/lmarena_ai/status/1891706264800936307 It's been said before but it is great news for consumers that there's so much competition in the LLM space. If it's hard for any one player to get daylight between them & the 2nd best alternative, hopefully that means one monopolistic firm isn't going to be sucking up all the value created by these things

It's not good news when this competition comes at cost of a gigantic over inflated bubble, in which all the big players keep on sucking billions from investors without even having a business model. This hype will burst sooner than later and will trigger yet another global recession. This is untenable.

The dot com bubble wiped out many billions of dollars in valuation.

The dot com bubble also gave us the most valuable companies in history, like Google, Apple, Amazon, Facebook, etc.

Re: Grok3 Launch [video]

#834
post #798

Earlier quoted context omitted.

It's a reflection of the wider society and (as others have pointed out) the media environment. HN can't be immune from macro trends. https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... We've been here before. It will likely subside, as past swings and fluctuations have. It always takes longer than it feels like it should, but in retrospect turns out to be shorter than it felt like it did.

haha interesting search query. thanks for your hard work dang!

(I feel bad about linking so often to my own comments but that information mostly doesn't exist anywhere else)

Re: Grok3 Launch [video]

#835
post #74

Earlier quoted context omitted.

I keep hearing about Claude's impressive coding skills (compared to its benches) yet, not evident for me (I use the web version, not cline). Compared to 4o it's not that great.

What are you using it for in general? IME the reason Claude pulls out ahead is that when you use it in a larger existing codebase, it keeps everything "in the style" of that codebase and doesn't veer off into weird territory like all the others.

My experience as well. Working in Scala primarily, it tends to be very good at following the constructs of the project.

Using a specific Monad-transformer regularly? It'll use that pattern, and often very well, handling all the wrapping and unwrapping needed to move data types about (at least well enough that the odd case it misses some wrapping/unwrapping is easy to spot and manage).

A custom GPT or GEM with the same source files, and those models regularly fail to maintain style and context, often suggesting solutions that might be fine in isolation but make little sense in the context of a larger codebase. It's almost like they never reliably refer to the code included in the project/GPT/GEM.

Claude on the other hand is so consistent about referring to existing artifacts that, as you approach the limit of project size (which is admittedly small) you can use up your entire 5-hour block of credits with just a few back-and-forths.

Re: Grok3 Launch [video]

#836

Earlier quoted context omitted.

Here's some things I have in my chatgpt history: - Discussing the various stages of candymaking and their relation to the fundamental properties of sugar syrups, and which candies are crystalline vs amorphous. It turns out junior mints are fudge. Fondant is really just fudge. Everything is fudge, my god. - Summarizing various SEC filings and related paperwork to understand the implications of an activist investor int…

> Everything is fudge, my god Fudge is made with milk - am I missing a joke?

Technically fudge is just a crystalline sugar candy with a certain water percentage. Milk is optional (and frequently omitted). Reese's peanut butter cups are fudge, for example.

Re: Grok3 Launch [video]

#837
post #490

Earlier quoted context omitted.

ChatGPT is literally generating billions in revenue. Cursor is the fastest growing company of all time. This lame HN trope of LLMs having no business model needs to die.

> ChatGPT is literally generating billions in revenue. It’s losing more billions than what it’s generating. Revenue does not equate profit. https://www.cnbc.com/2024/09/27/openai-sees-5-billion-loss-t...

They're still early on the growth curve where there's enough opportunity for future growth that investing in scaling and improvement is more important than turning an immediate profit.

Remember when everyone on HN was sure Uber would never be profitable? Or Tesla? Or Amazon?

Re: Grok3 Launch [video]

#838
Am I the only one who isn't impressed by this? Grok3 is failing basic OCR, react/sql coding excercises that Sonnet and Gemini completes successfully.

I'm also skeptical of lmarena as there is a large number of Elon Musk zealots trying to pass off Grok as a proxy for Tesla shares.

Re: Grok3 Launch [video]

#839

What are your first impressions using it? (Not available in Europe currently). Is it a game-changer?

No, it was underwhelming, failing basic coding tasks, OCR/Image recognitions that none of the other existing models screw up.

Re: Grok3 Launch [video]

#840

Earlier quoted context omitted.

The impression seems to be warranted: Grok 3 has directly jumpted to the top of all leaderboard categories in Chatbot Arena: https://lmarena.ai/?leaderboard In math it shares the top spot with o1 and is just a few points behind (well within errors). In creative writing it is basically ex-aequo with the latest ChatGPT 4o and in coding it's actually significantly ahead of everyone else and represents a new SOTA.

What do we do to assess the intelligence of these models after they are smarter than any human? From the kinds of questions it's answering seems like they are almost there. Do we have a way to tell if one model is smarter than another at that point?

Nah, at the end of the day "things that are easy for humans are [still] hard for computers, and vice versa". DeepBlue was super-human at chess and couldn't play tic tac toe. Today's AI is (almost?) super-human at math yet only very recently learned to play tic tac toe, and still can't learn to do anything - because it can't learn, and has no innate drives to expose itself to learning situations even if it could.

Here's a real world intelligence test. Take on each AI as a remote intern/new-hire, and try to train it to become a useful team member (solving math puzzles or manufacturing paperclips does not count).

Post reply on HN