Live data from Hacker News

Meta Llama 3

llama.meta.com

221–230 of 965 posts

Re: Meta Llama 3

#221

I’m impressed by the benchmarks but really intrigued by the press release with the example prompt ~”Tell me some concerts I can go to on Saturday”. Clearly they are able to add their Meta data to context, but are they also crawling the web? Could this be a surface to exfiltrate Meta data in ways that scraping/ APIs cannot?

It appears they're using Google for web searches, a la Perplexity.

Re: Meta Llama 3

#223
post #154

Earlier quoted context omitted.

Wild considering, GPT-4 is 1.8T.

Once benchmarks exist for a while, they become meaningless - even if it's not specifically training on the test set, actions (what used to be called "graduate student descent") end up optimizing new models towards overfitting on benchmark tasks.

"graduate student descent"

Ahhh that takes me back!

Re: Meta Llama 3

#224

Initial observations from the Meta Chat UI... 1. fast 2. less censored than other mainstream models 3. has current data, cites sources I asked about Trump's trial and it was happy to answer. It has info that is hours old --- Five jurors have been selected so far for the hush money case against former President Donald Trump ¹. Seven jurors were originally selected, but two were dismissed, one for concerns about her im…

It's likely RAG / augmented with web data. Would be interested if local execution returned the same results.

It is. You can see a little "G" icon indicating that it searched the web with Google.

Re: Meta Llama 3

#226

I am always excited to see these Open Weight models released, I think its very good for the ecosystem and definitely has its place in many situations. However since I use LLMs as a coding assistant (mostly via "rubber duck" debugging and new library exploration) I really don't want to use anything other than the absolutely best in class available now. That continues to be GPT4-turbo (or maybe Claude 3). Does anyone k…

Do you mind my asking, if you're working on private codebases, how you go about using GPT/Claude as a code assistant? I'm just removing IP and pasting into their website's chat interface. I feel like there's got to be something better out there but I don't really know anyone else that's using AI code assistance at all.

I built a desktop tool to help reduce the amount of copy-pasting and improve the output quality for coding using ChatGPT or Claude: https://prompt.16x.engineer/

Re: Meta Llama 3

#228

Earlier quoted context omitted.

Lex walked so that Dwarkesh could run. He runs the best AI podcast around right now, by a long shot.

I agree that it is the best AI podcast. I do have a few gripes though, which might just be from personal preference. A lot of the time the language used by both the host and the guests is unnecessarily obtuse. Also the host is biased towards being optimistic about LLMs leading to AGI, and so he doesn't probe guests deep enough about that, more than just asking something along the lines of "Do you think next token pre…

I struggle to blame people for speaking in whatever way is most natural to them, when they're answering hard questions off the cuff. "I apologize for such a long letter - I didn't have time to write a short one."

Re: Meta Llama 3

#229
post #16

Zuck has an interview out for it as well, https://twitter.com/dwarkesh_sp/status/1780990840179187715

Very interesting part around 5 mins in where Zuck says that they bought a shit ton of H100 GPUs a few years ago to build the recommendation engine for Reels to compete with TikTok (2x what they needed at the time, just to be safe), and now they are accidentally one of the very few companies out there with enough GPU capacity to train LLMs at this scale.

Re: Meta Llama 3

#230

Lots of great details in the blog: https://ai.meta.com/blog/meta-llama-3/ Looks like there's a 400B version coming up that will be much better than GPT-4 and Claude Opus too. Decentralization and OSS for the win!

The blog did not state what you said, sorry I’ll have to downvote your comment
Post reply on HN