Live data from Hacker News

LLaMA: A foundational, 65B-parameter large language model

ai.facebook.com

41–50 of 209 posts

Re: LLaMA: A foundational, 65B-parameter large language model

#41
Given the limited release of access under a license - is there any decision on whether a model like this would be protected by copyright if it leaked?

The output of an algorithm is not a human authored creative work, which given the recent decision in the book case would seem to weigh against it.

Re: LLaMA: A foundational, 65B-parameter large language model

#43
post #36

I was surprised to see this citation in the linked paper > Alan M Turing. 2009. Computing machinery and intelligence Did he invent a time machine too?

Looks like it was republished with additional commentary: https://link.springer.com/chapter/10.1007/978-1-4020-6710-5_...

Re: LLaMA: A foundational, 65B-parameter large language model

#44

Quick notes from first glance at paper https://research.facebook.com/publications/llama-open-and-ef... : * All variants were trained on 1T - 1.4T tokens; which is a good compared to their sizes based on the Chinchilla-metric. Code is 4.5% of the training data (similar to others). [Table 2] * They note the GPU hours as 82,432 (7B model) to 1,022,362 (65B model). [Table 15] GPU hour rates will vary, but let's give a ra…

https://github.com/facebookresearch/llama/blob/main/MODEL_CA... (linked in OP) has basic information about this model.

Re: LLaMA: A foundational, 65B-parameter large language model

#45
post #22

>'We release all our models to the research community' Where release means you fill out a form and wait indefinitely. Also no use for commercial purposes - which means 95% of users are out - certainly doesn't democratize LLMs.

I think these "open" models are the lamest release strategy. I get the idea of being closed à la OpenAI to protect a business model, and I understand open sharing in the community

"We'll let you see it if we approve you" is just being a self-important pompous middle man.

Re: LLaMA: A foundational, 65B-parameter large language model

#46
post #17

Earlier quoted context omitted.

Which kind of suggests Microsoft made a really bad move antagonizing the open-source community with Gtihub Copilot. They got a few years of lead time in the "AI codes for you" market, but in exchange permanently soured a significant fraction of their potential userbase who will turn to open-source alternatives soon anyway. I wonder if they'd have been better served focusing on selling Azure usage and released Copilot…

How did Microsoft sour developers with Copilot? I know dozens of people that pay for it (including myself) and I feel like it is widely regarded as a "no brainer" for the price that it's offered at. Please help me understand!

They didn’t. There is a small group of people that are always looking for the latest reason to be outraged and to point at any of the big tech companies and go “aha! They are evil!” Copilot’s ai was trained on GitHub projects and so these people are taking their turns clutching their pearls inside of their little bubble.

I’d bet that more than 95% of devs haven’t even heard of this “controversy” and even if they did, wouldn’t care.

Re: LLaMA: A foundational, 65B-parameter large language model

#47
post #37
post #11

Earlier quoted context omitted.

I don't see such grand stratagems as being a likely explanation. It seems more likely that a bunch of dorks are running around unsupervised trying their best to make lemonade with whatever budgets they are given as nobody can realistically manage AI researches. Or at least, it seems that much like in political governance, events are suddenly outpacing the ability of corporate to react. In the case of OpenAI it is a "…

> It seems more likely that a bunch of dorks are running around unsupervised trying their best to make lemonade with whatever budgets they are given Probably not, Zuck is announcing it.

  Today we're releasing a new state-of-the-art AI large language model called LLaMA designed to help researchers advance their work. LLMs have shown a lot of promise in generating text, having conversations, summarizing written material, and more complicated tasks like solving math theorems or predicting protein structures. Meta is committed to this open model of research and we'll make our new model available to the AI research community.
I don't know what that means or if he even wrote/read it tbh. I hope it literally just means Meta is actually committed to this open model of research (for now).

Maybe he is being a Machiavellian moat filler, I stand corrected. I think/hope that they don't really have a plan to counter OpenAI yet because I am afraid this attitude won't last once they do and this stuff has recently started moving quickly.

Re: LLaMA: A foundational, 65B-parameter large language model

#48
post #10

Earlier quoted context omitted.

> We release all our models to the research community. This is yet more evidence for the "AI isn't a competitive advantage" thesis. State-of-the-Art is a public resource, so competing with AI offers no "moat".

In terms of medieval warfare, what Facebook is doing here looks like filling the moat with rocks and dirt. OpenAI is worth billions, Microsoft is spending billions to retrofit most of their big offerings with AI, Google is doubtless also spending billions to integrate AI in their products. Their moat in all cases is “a blob of X billion weights trained on Y trillion tokens”. Facebook here is spending mere _millions_…

[deleted]

Re: LLaMA: A foundational, 65B-parameter large language model

#49
post #24
post #17

Earlier quoted context omitted.

How did Microsoft sour developers with Copilot? I know dozens of people that pay for it (including myself) and I feel like it is widely regarded as a "no brainer" for the price that it's offered at. Please help me understand!

The company that tried to kill Linux in the 90s, owned by the world's most famously rich man, is now stealing my code and selling it back to me? Yeah, fuck that.

I mean, that's like saying an author steals the open source alphabet and charges you for reading their ordering of letters, as if the ordering of letters isn't where all the value is.
Post reply on HN