Live data from Hacker News

Grok3 Launch [video]

x.com

451–460 of 1001 posts

Re: Grok3 Launch [video]

#451

[flagged]

I'm very sorry if this isn't the case, but this message really feels LLM-written.

Its because of the em dashes (- is a normal dash, — is an em dash). Very few real people use those outside of writing books or longform articles.

There's also some strange wordings like "back-pocket tests."

It's 100% LLM generated.

What is much scarier is that those "quick reply" blurbs on Android/Gmail (and iOS?) will be able to be trained on your entire e-mail and WhatsApp history. That model will have your writing mannerisms and even be a stochastic mimic of your reasoning. So, you won't be able to even realize a model answered you, not a real person. And the initial message the model is responding to might be written by the other person's personal model.

The future of digital interactions might have some sort of cryptographic signing guaranteeing you're talking to a human being, perhaps even with blocked copy-pasting (or well, that part of the text shows up as unverified) and cheat detection.

Going even a layer deeper / more meta: what does it ultimately matter? We humans yearn for connection, but for some reason that connection only feels genuine with another human. Whereas, what is the difference between a human typing a message to you, a human inhabiting a robot body, a model typing a message to you, and a model inhabiting a robot body, if they can all give you unique interactions?

Re: Grok3 Launch [video]

#452
post #10

Grok has gotten to the top of one benchmark: https://x.com/lmarena_ai/status/1891706264800936307 It's been said before but it is great news for consumers that there's so much competition in the LLM space. If it's hard for any one player to get daylight between them & the 2nd best alternative, hopefully that means one monopolistic firm isn't going to be sucking up all the value created by these things

Probably bad news for the vendors, though. I genuinely struggle to see how most of these LLM companies are going to monetize and profit off their efforts with LLMs already in commodity territory. Government contracts can only flow for so long?

Government contracts are so big a few of them can sustain a F500 company; for AI, many CDAO contracts are 50-500MM$. If they do a big SI project with it, could be 1-2B$. Money is also guaranteed over 5 years and if the program doesn't get shuttered, the contract will renew at that point (or go to recompete).

That being said it's my understanding that these companies don't have many huge contracts at all -- you can audit this in like 10 minutes on FPDS. Companies need a LOT of capital, time, and expertise to break into the industry and just compliance audit timelines are 1-4 years right now, so this could definitely change in the next couple years.

Re: Grok3 Launch [video]

#453

Off topic, but just in case: is there a good reference on how people actually use LLMs on a daily basis ? All my attempts so far have been pretty underwhelming: * when I use chatbots as search engines, I'm very quickly disappointed by obvious hallucinations * I ended up disabling github copilot because it was just "auto-complete on steroids" at best, and "auto-complete on mushrooms" at worst * I rarely have use cases…

The only plausible explanation for the amount of resources poured into these language models is the hope that they somehow become the origin of AGI, which I think is pretty fanciful. I can feel the cold wind of the next AI winter coming on. It's inevitable. Computers are good at emulating intelligent behavior, people get excited that it's around the corner, and the hype boils over. This isn't the last time this will…

I predict this comment will age very, very poorly. Bookmarked.

Re: Grok3 Launch [video]

#454

Off topic, but just in case: is there a good reference on how people actually use LLMs on a daily basis ? All my attempts so far have been pretty underwhelming: * when I use chatbots as search engines, I'm very quickly disappointed by obvious hallucinations * I ended up disabling github copilot because it was just "auto-complete on steroids" at best, and "auto-complete on mushrooms" at worst * I rarely have use cases…

> I ended up disabling github copilot because it was just "auto-complete on steroids" at best

this is good enough sell for me, and it's like sub 1-in-50 that it's "auto-complete on mushrooms" (again my experience, YMMV).

An awful lot of the time, my day to day work involves writing one piece of code and then copy-pasting it changing a few variable names. Even if I factor out the code into a method, I've still got to call that method with the different names. CoPilot takes care of that drudgery and saves me countless minutes per day. It therefore pays for itself.

I also use ChatGPT every time I need some BASH script written to automate a boring process. I could spend 20-30 minutes searching for all the commands and arguments I would need, another 10 minutes typing in the script, another 10-20 minutes debugging my inevitable mistakes. Or I make sure to describe my requirements exactly (5-10 minutes), spend 5 minutes reviewing the output, iterate if necessary (usually because I wasn't clear enough in the instructions).

3-5x speed up for free. Who's not going to take that win?

Re: Grok3 Launch [video]

#456
post #10

Grok has gotten to the top of one benchmark: https://x.com/lmarena_ai/status/1891706264800936307 It's been said before but it is great news for consumers that there's so much competition in the LLM space. If it's hard for any one player to get daylight between them & the 2nd best alternative, hopefully that means one monopolistic firm isn't going to be sucking up all the value created by these things

Probably bad news for the vendors, though. I genuinely struggle to see how most of these LLM companies are going to monetize and profit off their efforts with LLMs already in commodity territory. Government contracts can only flow for so long?

Like always through ads.

Re: Grok3 Launch [video]

#457
post #438

Earlier quoted context omitted.

Probably bad news for the vendors, though. I genuinely struggle to see how most of these LLM companies are going to monetize and profit off their efforts with LLMs already in commodity territory. Government contracts can only flow for so long?

Yes, I couldn't have imagined that in the end, the AI wrappers are where the money is.

What if the money isn't there either? What if this AI thing lowers costs of everything it touches without generating meaningful financial returns itself?

Re: Grok3 Launch [video]

#458

Karpathy gave his initial impression: https://x.com/karpathy/status/1891720635363254772 The pull quote is: The impression overall I got here is that this is somewhere around (OpenAI) o1-pro capability

Naive question from a bystander , but since DeepSeek is open source and is on par with o1-pro (is it?), shouldn't we expect that anybody with the computer power is capable to compete with o1-pro?

You'd still need a fairly large amount of compute power to be able to run DeepSeek R1 locally, no?

Re: Grok3 Launch [video]

#459

Karpathy gave his initial impression: https://x.com/karpathy/status/1891720635363254772 The pull quote is: The impression overall I got here is that this is somewhere around (OpenAI) o1-pro capability

Naive question from a bystander , but since DeepSeek is open source and is on par with o1-pro (is it?), shouldn't we expect that anybody with the computer power is capable to compete with o1-pro?

> DeepSeek is open source and is on par with o1-pro (is it?)

There is no being "on par" in this space. Model providers are still mostly optimising for a handful of benchmarks / goals, like we can already see that Grok 3 is doing incredibly well on human preference (LM Arena) however with Style Control, it's suddenly behind ChatGPT-4o-latest and Gemini 2.0 is out the picture. So even within a single domain, goal, benchmark—it's not as straightforward as to say that one model is "on par" with another.

> shouldn't we expect that anybody with the computer power is capable to compete with o1-pro?

Not necessarily. I know it may be tempting to think that Grok 3 is entirely a result of xAI having lots of "computer power", but you have to recognise that this mindset is coming from a place of ignorance, not wisdom. Moreover, it doesn't even pass off as "cynical" view, because it's common knowledge that model training is really, really complicated. DeepSeek results are note-worthy, and really influential in some respects, but it hasn't magically "solved" training, or made training necessarily easier / less expensive for the interested parties. They never shared the low-level performance improvements, just model weights and lots of insight. For talented researchers, this is valuable, of course, but it's not like "anybody" could easily benefit from it in their training regimes.

Update: RFT (contra SFT) is becoming really popular with service providers, and it's not been "standardised" beyond whatever reproductions to have emerged in the weeks prior, moreover R1 cost is still pretty high[1] at something like $7/Mtok, & bandwidth is really not great. Consider something like Google Vertex AI's batch pricing for Gemini 1.5 Pro and Gemini 2.0 Flash models, which is at 50% discount, and their prompt caching which is at 75% discount. R1 is still got a way to go.

[1]: https://openrouter.ai/deepseek/deepseek-r1/providers?sort=th...

Re: Grok3 Launch [video]

#460
post #74

Earlier quoted context omitted.

And Anthropic not even in the top 10 ...

I keep hearing about Claude's impressive coding skills (compared to its benches) yet, not evident for me (I use the web version, not cline). Compared to 4o it's not that great.

I spend four to five hours coding per day and subscribe to every major LLM and Claude is still by far the best for me personally and my co workers.
Post reply on HN