Live data from Hacker News

Grok3 Launch [video]

x.com

161–170 of 1001 posts

Re: Grok3 Launch [video]

#161

https://garymarcus.substack.com/p/elon-musks-terrifying-visi... I'm not sure if this was a very bad joke by Elon, or if Grok 3 is really biased like that.

I am not sure why people pay attention to Gary Marcus. He isn’t an expert in AI. And if you followed him in the past at all, it is obvious he has a huge amount of political bias. It is really telling that he repeatedly goes after Elon Musk, and is now making bizarre unfounded claims about propaganda, but didn’t have nearly as much to complain about with DeepSeek, which has literal government propaganda.

Re: Grok3 Launch [video]

#162

https://garymarcus.substack.com/p/elon-musks-terrifying-visi... I'm not sure if this was a very bad joke by Elon, or if Grok 3 is really biased like that.

Karpathy, which is IMHO a serious and balanced person, lamented that it looks too censored (see recent tweets). Elon Musk is (for me) a very scary person, and it is important to evaluate AI safety (but I believe that the safety that matters in AI is of a different kind), yet to listen to Gary Marcus does not make any sense: it's just an extremely biased person that is riding the anti AI wave.

Yes that's certainly true. I was a bit hesitant to post a link from Gary Marcus. But I was mostly posting it for the Elon tweet. I assume the tweet is not fake. So you can ignore about Garys opinion here and just take Elons tweet as it is.

Re: Grok3 Launch [video]

#165

[flagged]

Because he and his organization have demonstrated ignorance of the services he's not only auditing, but making pretty substantial cuts to. One example I'm familiar with, cutting up to 10% of the personnel to the Technology Transformation Services at GSA is quite likely to reduce the efficiency of both government and private sector government contractors.

https://news.ycombinator.com/item?id=43037624

Re: Grok3 Launch [video]

#167

Companies have hijacked the open source concept to mean downloadable blob and we follow them as I see in the comments. It’s a real shame.

An llm isn't software any more than a matrix is.

What do you think an open source matrix should look like?

Re: Grok3 Launch [video]

#169
Off topic, but just in case: is there a good reference on how people actually use LLMs on a daily basis ? All my attempts so far have been pretty underwhelming:

* when I use chatbots as search engines, I'm very quickly disappointed by obvious hallucinations

* I ended up disabling github copilot because it was just "auto-complete on steroids" at best, and "auto-complete on mushrooms" at worst

* I rarely have use cases where I have to "generate a plausible page of text that statistically looks like the internet" - usually, when I have to write about something, it's to put information that's in my head into other people head

* I'd love to have something that reads all my codebase and draws graphs, explain how things work, etc... But I tried aider/ollama, etc.. and nothing even starts making sense (is that an avenue to persevere in, though ?)

* At once, I tried to write in plain english a situation where a team has to do X tasks, in Y weeks, and I needed a table of who should be working on what for each week. I was impressed that LLMs were able to produce a table - the slight problem was that, of course, the table was completely wrong. Again, is it just bad prompting ?

It's an interesting problem when you don't know if you're just having a solution in search of a problem, or if you're missing something obvious about how to use a tool.

Also, all introductory texts about LLMs go into many details about how they're made (NNs and transformers and large corpuses and lots of electricity etc...) but "what you can do with it" looks like toy examples / simply not what I do."

So, what is the "start from here" about what it can really do ?

Re: Grok3 Launch [video]

#170
post #10

Grok has gotten to the top of one benchmark: https://x.com/lmarena_ai/status/1891706264800936307 It's been said before but it is great news for consumers that there's so much competition in the LLM space. If it's hard for any one player to get daylight between them & the 2nd best alternative, hopefully that means one monopolistic firm isn't going to be sucking up all the value created by these things

Who cares about benchmarks?

These things still cost me time because of hallucinations.

Post reply on HN