Live data from Hacker News

Grok3 Launch [video]

x.com

991–1000 of 1001 posts

Re: Grok3 Launch [video]

#992

Earlier quoted context omitted.

This is akin to suggesting that we should have all been praising Microsoft for their achievements back in the day rather than saying a word about EEE, their monopolism, or their enmity towards open source. Or that it’s not polite to bring up the CCP when discussing TikTok. Bottom line: a technology that has the ability to shape human thought perhaps more than any other in history is owned by a man with some truly vil…

[flagged]

I don't think anyone is telling you what your opinions should be. The GP post just presents the GP's opinion. You're free to agree or disagree with it as you choose.

If you read a comment that you're unhappy with, downvote it and move on.

Re: Grok3 Launch [video]

#994
post #153

Karpathy believes that this is at o1-pro level[1]. This again proves that OpenAI simply has no tech moat whatsoever. Elon's $97 billion offer for OpenAI last week was reasonable given that xAI already have something just a few months behind - it would probably be faster for xAI to catch up with o3 than going through all those paperworks and lawyer talks required for such an acquisition. Elon also has some huge up-han…

Firstly, the 97Bn was for the non-profit, not for the company. The company is being valued in funding rounds closer to 300Bn. I think it may be true that OpenAI has no moat, but if it has no moat then all of these AI companies are overvalued (including xAI) and Elon should just stop bothering to throw his money at it. I would say Elon probably actually doesn't have much of an advantage here. In both SpaceX and Tesla…

> SpaceX consumed enormous amounts of cash

No, spacex projects are extremely $ efficient. The total project cost of starship is like 20% of the SLS.

> he's competing against all the biggest companies in the world in this race.

No, this is a not a pissing contest on who has the most $. If it is about who can come up with most $, then the entire race is already over as the CCP has access to trillions of $ CASH.

Re: Grok3 Launch [video]

#995
post #74

Earlier quoted context omitted.

I keep hearing about Claude's impressive coding skills (compared to its benches) yet, not evident for me (I use the web version, not cline). Compared to 4o it's not that great.

I prototyped on the weekend and started out with 4o because i had a subscription running. After an hour and a half assed working result, i put everything into claude and it made it significant better on the first try and i had not a subscription active with claude.

Really interesting, I used it today still lots of issues. Maybe my python notebook is not approach is too complicated for Sonnet? Couldn't be able to fix a custom complex seaborn plot. 4o failed too. o3-mini-high managed to do it really well on the other hand.

Re: Grok3 Launch [video]

#997

Earlier quoted context omitted.

ChatGPT is literally generating billions in revenue. Cursor is the fastest growing company of all time. This lame HN trope of LLMs having no business model needs to die.

What source(s) are there for cursor's growth rate/revenue ?

So, answering my own question, there is this.

https://sacra.com/research/cursor-at-100m-arr/

Sounds legit.

Re: Grok3 Launch [video]

#999

Earlier quoted context omitted.

The impression seems to be warranted: Grok 3 has directly jumpted to the top of all leaderboard categories in Chatbot Arena: https://lmarena.ai/?leaderboard In math it shares the top spot with o1 and is just a few points behind (well within errors). In creative writing it is basically ex-aequo with the latest ChatGPT 4o and in coding it's actually significantly ahead of everyone else and represents a new SOTA.

lmarena/lmsys is beyond useless, looking at prior rankings of models vs formal benchmarks or testing for accuracy + correctness on batches of real world data. It's a bit like using a poll of Fox News to discern the opinions of every American; the audience voting is consistently found wanting. Not even getting into how easily a bad actor with means + motivation (in this "hypothetical" instance wanting to show that a c…

So you're saying that either A: users interacting with models can't objectively rate what responses seem better to humans, B: xAi as a newcomer has somehow managed to game the leaderboard better than all those other companies, or C: all those other companies are not doing it. By those standards every test ever devised for anything is beyond useless. But simply not having the model creator running the evaluation is already going a long way.

Re: Grok3 Launch [video]

#1000
post #510
post #478

Earlier quoted context omitted.

The full-sized DeepSeek-R1 is on par with o1. o1-pro is "o1 on steroids" and was the first selling point of the $200/month Pro subscription but they later also added "Deep Research" and Operator to the Pro subscription.

Every year seems like we get worse at naming things in non confusing ways. I am waiting for the o1-pro-max now, pro max ultra and pro max ultra plus.

Off by one and naming things.
Post reply on HN