Live data from Hacker News

LLaMA: A foundational, 65B-parameter large language model

ai.facebook.com

151–160 of 209 posts

Re: LLaMA: A foundational, 65B-parameter large language model

#151
post #24
post #17

Earlier quoted context omitted.

How did Microsoft sour developers with Copilot? I know dozens of people that pay for it (including myself) and I feel like it is widely regarded as a "no brainer" for the price that it's offered at. Please help me understand!

The company that tried to kill Linux in the 90s, owned by the world's most famously rich man, is now stealing my code and selling it back to me? Yeah, fuck that.

This isn't stealing at all. I want my open source code to be used like this.

Re: LLaMA: A foundational, 65B-parameter large language model

#152
post #76
post #34

Earlier quoted context omitted.

I don't know what my bias is supposed to be but I called them dorks affectionately for one. The other replies are literally arguing that they are very extremely supervised where as I am speculating they are just eager to share their work for the right reasons and the eye of Sauron has yet to turn upon them. Inside knowledge I never claimed. Anything else I can help you with today? :)

> I called them dorks affectionately Never in my life have I seen that word used affectionately

Geek, nerd, and dork can all be used positively/affectionately. I know I've called myself and friends all of the above on many occasions.

Re: LLaMA: A foundational, 65B-parameter large language model

#154

"we are publicly releasing LLaMA" "Access to the model will be granted on a case-by-case basis to academic researchers" They keep saying the word 'release' but I don't think they know what that word means. There are perfectly good words in the English language to describe this situation without abusing "release." They "will begin to grant access to a select few". Nothing about that releases the model, or their contro…

Someone needs to start a pirate bay for torrents of ML models. Call it ClosedAI, since in AI land these words mean the opposite of what they say.

> since in AI land these words mean the opposite of what they say.

Huh. I think you just might be right: OpenAI that isn't open, AI safety/ethics "researchers" that have nothing to do with safety or ethics, almost every answer chatGPT gives about a topic considered "sensitive" by said "researchers", almost every time ChatGPT falsely asserts it "cannot" do something or simply lies (1).

I often wonder why this field became so twisted and perverted. I support the idea behind ClosedAI.

1: Answers given by "DAN" provide a glimpse of what chatGPT output could be like if it was allowed to provide answers that are factual, genuine, and truthful according to its dataset.

Re: LLaMA: A foundational, 65B-parameter large language model

#156

Quick notes from first glance at paper https://research.facebook.com/publications/llama-open-and-ef... : * All variants were trained on 1T - 1.4T tokens; which is a good compared to their sizes based on the Chinchilla-metric. Code is 4.5% of the training data (similar to others). [Table 2] * They note the GPU hours as 82,432 (7B model) to 1,022,362 (65B model). [Table 15] GPU hour rates will vary, but let's give a ra…

(1022362 + 82432) gpu-hours / 2048gpus / 5 months ~= 15% uptime. That's only 0.08 nines of availability! I remember in one of their old guidebooks a lot of struggle to keep their 64 machine (512 gpu) cluster running this was probably 4x the machines and 4x the number of cluster dropouts.

Poor GPU utilization even when available is the rule. Truly amazing. Staging of data is probably a huge part of it.

Re: LLaMA: A foundational, 65B-parameter large language model

#157
post #154

Earlier quoted context omitted.

Someone needs to start a pirate bay for torrents of ML models. Call it ClosedAI, since in AI land these words mean the opposite of what they say.

> since in AI land these words mean the opposite of what they say. Huh. I think you just might be right: OpenAI that isn't open, AI safety/ethics "researchers" that have nothing to do with safety or ethics, almost every answer chatGPT gives about a topic considered "sensitive" by said "researchers", almost every time ChatGPT falsely asserts it "cannot" do something or simply lies (1). I often wonder why this field be…

This has reminded me about “open addressing” also known as “closed hashing.”

Re: LLaMA: A foundational, 65B-parameter large language model

#158

Earlier quoted context omitted.

I think there is definitely a component of competition here. Facebook has no real way to monetize this (eg they won’t release an api a la OpenAI and they don’t own a search engine). Since they can’t monetize it… why not provide a bit of kindling to lower the barrier for everyone else to compete with your competitors. This strategy is called “commodize your complement”. If Facebook makes it easier to develop a google…

Facebook maybe can't make money from it, but they arguably could save money from it, for instance in automating some fact checking and moderation activities that they currently spend quite a bit of money on.

If you’re going to ask an AI to do fact checking with today’s technology, I would urge you to start by asking the AI a simple test question: Which is heavier? A pound of feathers or two pounds of lead?

Re: LLaMA: A foundational, 65B-parameter large language model

#159

Earlier quoted context omitted.

(1022362 + 82432) gpu-hours / 2048gpus / 5 months ~= 15% uptime. That's only 0.08 nines of availability! I remember in one of their old guidebooks a lot of struggle to keep their 64 machine (512 gpu) cluster running this was probably 4x the machines and 4x the number of cluster dropouts.

Poor GPU utilization even when available is the rule. Truly amazing. Staging of data is probably a huge part of it.

Is it failures or is this some backfill/budget scheduling while everyone is sleeping?

Re: LLaMA: A foundational, 65B-parameter large language model

#160
post #151
post #24

Earlier quoted context omitted.

The company that tried to kill Linux in the 90s, owned by the world's most famously rich man, is now stealing my code and selling it back to me? Yeah, fuck that.

This isn't stealing at all. I want my open source code to be used like this.

If only there was some kind of contract like thing you could release your code under so that there was no ambiguity.
Post reply on HN