Live data from Hacker News

Zuckerberg disses closed-source AI competitors as trying to 'create God'

techcrunch.com

41–50 of 64 posts

Re: Zuckerberg disses closed-source AI competitors as trying to 'create God'

#43

Earlier quoted context omitted.

just putting weights on hug is not enough for a model to be open. it is like sharing the binaries of a software. they should put their entire code on github and show how it was trained, architecture and everything else. also, the most important thing about a model is not it's architecture, but the dataset it was trained on. that is the "it" in a model.

I am not sure why people keep pretending like this is possible or keep asking it to happen. You know that distributing all the data would be illegal copyright infringement, right? Also, its not like they could distribute the (petabytes?) worth of data to every rando who asks. Do you just not want open weight models to exist, and you are using a fake call for it to be more open, so as to get this stuff all shut down,…

Perhaps a good solution would be to form a trust responsible for storing and distributing training data. It could be held in common by all contributors and allow anyone to access high quality curated data. It would need legal allowances and exemption from copyright which would be difficult to work out but perhaps some mechanism could be created where copyright holders could derive benefit from not seeking to have their works removed, not necessarily monetary. This would at least get data collection out of the shadows and allow companies and researchers to attest that they are using the data only for training and not distributing it otherwise. That could go a long way toward clarifying issues of whether or not a book or article was in a dataset, how often it appeared, how many epochs it was used in, whether it only appeared as referenced or quoted by other text, etc. With that kind of disclosure it would be possible to write tests to ensure that the models won’t simply regurgitate the material without attribution or unprompted and publish those tests to assuage fears.

Re: Zuckerberg disses closed-source AI competitors as trying to 'create God'

#45
post #19

I would agree. AI is the machine god realized. Not by virtue of its capabilities, but by the agency people hand off to it. Whether it be creative tasks or just planning their activities, I see more people rely on AI for more decisions. They don’t want to think about something, so they have an AI choose for them. One could argue Top 10 lists and Guides already did this, but for those it was limited to certain aspects…

[dead]

Re: Zuckerberg disses closed-source AI competitors as trying to 'create God'

#46
The analytics and insights they'll generate alone, citing most of these Ai tools are harvesting based on Google accounts, will be staggering in hindsight. By using Ai tools, so many people are surrendering volumes of data that are so much more invasive than Social Media ever could be. As Chat GPT gets integrated into many services, they'll be enabled to build a far deeper profile on everyone that registers, and even a lot of data on people that don't register on the service.

I think we need to establish prominent technical leadership/chief advisor positions within the White House and each chamber of Congress now to be able to establish ethical baselines and oversight before tech managed by private companies turns even more predatory upon public data (Similar to The Surgeon General's Role in authority perhaps)... There should already be solid ground rules for operation and oversight of Ai development efforts in place to control the egos of the people that will have access to all of this user data, those types almost always exhibit God complexes without any real accountability nor responsibility, because most of the people that legislate have no clue of what they can see on the backend.

For example alone, if we though insider training was bad, just imagine how toxic it's going to get when Ai development employees, or a product owner has access to every aspect of social, economic, and business moves of every single company and individual that uses tools integrated across Chat GPT and similar services, including Email and MS Office (without anyone really knowing their data is easily accessible)... It's not a bright future for fairness and equity at all, and there will be far more oligarchs and monopolies than there are now.

Re: Zuckerberg disses closed-source AI competitors as trying to 'create God'

#47
post #19

I would agree. AI is the machine god realized. Not by virtue of its capabilities, but by the agency people hand off to it. Whether it be creative tasks or just planning their activities, I see more people rely on AI for more decisions. They don’t want to think about something, so they have an AI choose for them. One could argue Top 10 lists and Guides already did this, but for those it was limited to certain aspects…

> LLMs in particular are ready and eager to make a choice for you about almost anything. Wat. They're still just highly advanced auto-correct, and aren't eager to do anything. Sure some users might be eager to delegate the decision to them...

Of course they don't have real agency. Its a figure of speech.

Re: Zuckerberg disses closed-source AI competitors as trying to 'create God'

#49

[flagged]

If it flops we'll have loads of cheap GPU's :)

At this point, I think Nvidia has enough of an incentive to chase AI chips and their specific needs as an entirely different product. But yeah, for home users who give up, definitely.

Re: Zuckerberg disses closed-source AI competitors as trying to 'create God'

#50
post #35

Earlier quoted context omitted.

just putting weights on hug is not enough for a model to be open. it is like sharing the binaries of a software. they should put their entire code on github and show how it was trained, architecture and everything else. also, the most important thing about a model is not it's architecture, but the dataset it was trained on. that is the "it" in a model.

I'm not in the game, so pardon my ignorance, but is there any company that released the dataset the model was trained on for LLMs? To my understanding, it will lead to never-ending lawsuits due to grey areas as soon as they're released, and no company actually wants to go through that pain.

StableDiffusion 1.0 was trained on LAION-5B Dataset, but the code used to train it wasn't released, so even then it doesn't qualify as open source.
Post reply on HN