Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

481–490 of 648 posts

Re: GPT-4 details leaked?

#481

Earlier quoted context omitted.

Capitalism is not free trade between two parties. That's free market. Capitalism is control by capital.

There is no free market without Capitalism. Capitalism is where the private citizens controls the capital. Not the government. It's not possible to have a free market without competitive markets, price systems, private property, property rights recognition, and voluntary exchange.

??? Bazaars existed for millennia before capitalism arose.

> Capitalism is where the private citizens controls the capital.

Well, that's a tautology, but based on the examples you cite in the next sentence, you're mixing up "property" with "capital". Capital is specifically the means of production. All capital is property, but not all property is capital. A factory filled with machinery is capital; your toothbrush isn't.

> It's not possible to have a free market without competitive markets, price systems, private property, property rights recognition, and voluntary exchange.

Most of these predate capitalism.

Re: GPT-4 details leaked?

#483

Earlier quoted context omitted.

{Hypothesis, test, loop} is the scientific method, and I can guarantee it is being used when fine tuning an LLM.

That's a common interpretation of what science is, but it largely ends up being driven by confirmation bias. Because how do you know your hypothesis and test are even really connected? Or what you're seeing is a cause and not a correlation? This is why you need arguably the two most important factors in "real" science: predictability and falsifiability. Predictability means that if your hypothesis is correct then you…

We've got lots of datasets at this point in history - it ends up being pretty reasonable to formulate a hypothesis, propose a change to the ML system, then see how it does on a host of datasets. Most papers focus on one or two datasets, getting some initial improvement going from train-to-test, but the successful ideas end up transferring to a wide range of contexts. The extension to other dataseets + contexts constitutes the prediction of performance on other systems.

And sometimes (often?) techniques do fail to transfer! IME, it's only the simplest ideas which do transfer to new contexts well - there's a lot of dark heat in pushing Imagenet another hundredth of a point, which you realize when you try to take incremental 'new SOTA' papers and apply them to audio problems.

But the things which /do/ transfer constitute real advances. And that's (drumroll) science! Sometimes things look good initially and get falsified. Medicine is still science, even if drugs get disqualified in phase iii trials...

Re: GPT-4 details leaked?

#484

Earlier quoted context omitted.

Maybe because we're on the verge of being able to create fires which can actually consume the only home we have? Playing with fire is in large part an ego and greed issue. Yes, it allows us to dominate, but at what cost? I'd rather live a more balanced life than a greedy and ego driven life. I may not own the world, but I can be happy and sleep sound at night, and that matters.

We had nuclear weapons for almost 80 years and the world still hasn't ended. And I think that nuclear weapons are way more dangerous than Markov chains on steroids.

What if Wargames had LLMs involved?

Re: GPT-4 details leaked?

#485
post #440

Earlier quoted context omitted.

{Hypothesis, test, loop} is the scientific method, and I can guarantee it is being used when fine tuning an LLM.

The reason that this is Engineering as opposed to Science, is that the hypothesis is just, "hey, maybe this will work". Nobody has really explained why it works.

Ask Feynman why magnets works... And he'll say no one knows. https://www.youtube.com/watch?v=36GT2zI8lVA

It's the same in ML. Good predictions do come from real ideas about how the systems work - ideas in information theory, entropy, cognitive science, statistical mechanics and so on - formulated in a context of our existing understanding and prior results. That's science.

Re: GPT-4 details leaked?

#487

Earlier quoted context omitted.

Some kind of SETI project, but for training a high number parameter llm would be awesome.

How many people have A100s at home?

The comparison to SETI@home[1] is that you just need the same mount of total processing power, not the same supercomputer setup.

People don't need to own A100s, they just need to be willing to be part of a distributed supercomputer by running a background app that downloads chunks of data, processes them, and sends the result back. The utility comes from having enough people participate (which worked quite well for SETI@home, but helping find "signals from outer space" is a little bit more interesting than "helping train an LLM")

[1] https://setiathome.berkeley.edu/

Re: GPT-4 details leaked?

#488

Earlier quoted context omitted.

If anyone has a copyright claim to an LLM, the creators of the input data have more of a copyright claim than the company that trained it. There's a good chance they are not copyrightable at all. I'd bet there's a lot of people willing to take on that risk. However, they might still fall under trade secret law.

Why would an LLM be any less copyrightable than any other piece of software?

The "software" part of an LLM is pretty trivial -- the interesting piece is the the weights. Since the weights are mechanically generated by a computer, it can be argued that the weights are not copyrightable, just like a photograph taken by a monkey isn't copyrightable.

Re: GPT-4 details leaked?

#489

"Open" AI, a charity to benefit us all by pushing and publishing the frontier of scientific knowledge. Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. https://github.com/ggerganov/llama.cpp https://github.com/openlm-research/open_llama https://huggingface.co/TheBlok…

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science. Machine learning, as it is practiced to…

They are doing computer science.

Re: GPT-4 details leaked?

#490

Earlier quoted context omitted.

I'm tired of science as a religion. People treat it as some gospel, like if you check out some criterions you're suddenly "scientific" and instantly get a sense of validity and authority that you shouldn't logically get. I judge things as "what you can do", not "what can you predict". The only demonstration of knowledge and understanding is being able to do something. Not predict. Not "scientific method" and ridiculo…

Science is not academia. You say you prefer to judge things as "what you can do". Well, how do you judge "what you can do"? Astrologists, homeopaths, podiatrists, Christian scientists (!!!) and other such "heretics and mad men" rejected by the scientific establishment, will all tell you that they "can do" stuff, and so will all their many paying customers. How do we know they can't do what they say? Because science g…

And whence the tools to judge that

> science gives you the tools to know that you're wrong.

Post reply on HN