Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

471–480 of 648 posts

Re: GPT-4 details leaked?

#471

Earlier quoted context omitted.

Right. In historical context the scientific method was a huge improvement over the previous method of understanding the world, which was mostly religion. But today, academia is a victim of organizational and political capture, making it less competitive for talent.

People struggle to differentiate the scientific method and the lifecycle of academia.

In some ways academia is akin to a religion or cult that sprang up around the scientific method, and the cult has now diverged as far from the seed as far right christianity has from jesus’s teachings

Re: GPT-4 details leaked?

#472

Earlier quoted context omitted.

If anyone has a copyright claim to an LLM, the creators of the input data have more of a copyright claim than the company that trained it. There's a good chance they are not copyrightable at all. I'd bet there's a lot of people willing to take on that risk. However, they might still fall under trade secret law.

Why would an LLM be any less copyrightable than any other piece of software?

The model weights could be seen as a derived work, for which they didn't get the permission of the original copyright holders. Alternatively, it can be argued that the LLMs are no different than a fanfic writer trying to imitate the style of their favor author.

It's not obvious which way it will go, but I can see the point of those arguing that LLM data are ill-gotten gains.

Re: GPT-4 details leaked?

#473

Earlier quoted context omitted.

> The google memo was right about the lack of a moat. 5 months on, and nobody has yet beaten their result quality. I think there is a moat. Also, I think for many usecases, smarter is better. If a few cents can buy a more accurate answer, then it is always worth paying those few cents. So, while more hardware and more data can train a bigger better model, then that is the moat.

The moat is there until someone releases (or leaks) comparable training data.

And that gets more difficult every day, as previously accesible sources of data turn off their api.

though google may have something up its sleeve with the corpus of google books! I have been wondering if openAI secretly pulled in scihub or zlibrary to neutralize that potential advantage.

Re: GPT-4 details leaked?

#474
> Part of this extremely low utilization is due to an absurd number of failures requiring checkpoints that needed to be restarted from.

Hahahaha, the truth of anyone who has worked with quanty types running Python code at scale on a cluster

Re: GPT-4 details leaked?

#475

"Open" AI, a charity to benefit us all by pushing and publishing the frontier of scientific knowledge. Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. https://github.com/ggerganov/llama.cpp https://github.com/openlm-research/open_llama https://huggingface.co/TheBlok…

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science. Machine learning, as it is practiced to…

> Machine learning, as it is practiced today, is not science. There is no scientific theory behind it and there is no scientific method applied. There are no scientific questions asked, or attempted to be answered.

Total horseshit. There are tons of scientific papers on ML published. In fact it is MORE like traditional science than typical CS, because it is trying to reverse engineer how something we encountered in the real world works. We know NNs do amazing things, and we don't fully yet understand how.

Re: GPT-4 details leaked?

#476

Earlier quoted context omitted.

> as in not giving out dangerous answers that get people killed Literally taken, that is quite close to impossible. It was news days ago of somebody who committed suicide after having an interaction with a bot about nuclear risks or similar. To avoid that, the bot would have to be a high-ranking professional psychologist with an explicit purpose not to trigger destructive reactions. And that would fail the nature of…

"I can't conceive of such failure mode thus we're safe"

Sorry, I do not understand what you mean...

Re: GPT-4 details leaked?

#477

Earlier quoted context omitted.

I'm tired of science as a religion. People treat it as some gospel, like if you check out some criterions you're suddenly "scientific" and instantly get a sense of validity and authority that you shouldn't logically get. I judge things as "what you can do", not "what can you predict". The only demonstration of knowledge and understanding is being able to do something. Not predict. Not "scientific method" and ridiculo…

Just add epicycles.

Add enough epicycles and you've got a Fourier transform... Epicycles were a great idea but applied for the wrong reason.

Re: GPT-4 details leaked?

#478

Earlier quoted context omitted.

{Hypothesis, test, loop} is the scientific method, and I can guarantee it is being used when fine tuning an LLM.

That's a common interpretation of what science is, but it largely ends up being driven by confirmation bias. Because how do you know your hypothesis and test are even really connected? Or what you're seeing is a cause and not a correlation? This is why you need arguably the two most important factors in "real" science: predictability and falsifiability. Predictability means that if your hypothesis is correct then you…

ML has both of those things. In fact the situation is better than in many other sciences because the evaluation is completely digital. When you run an ML experiment you get back an objective measure in minutes or hours telling you if the model performed better or worse. If it's worse, boom, falsified. Many current papers are making predictive testable claims -- that based on various things we know what ideal parameter values are, the types of problems current NN archs will perform well on or poorly on, etc.

Re: GPT-4 details leaked?

#479

Earlier quoted context omitted.

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science. Machine learning, as it is practiced to…

I'm tired of science as a religion. People treat it as some gospel, like if you check out some criterions you're suddenly "scientific" and instantly get a sense of validity and authority that you shouldn't logically get. I judge things as "what you can do", not "what can you predict". The only demonstration of knowledge and understanding is being able to do something. Not predict. Not "scientific method" and ridiculo…

Science is not academia. You say you prefer to judge things as "what you can do". Well, how do you judge "what you can do"? Astrologists, homeopaths, podiatrists, Christian scientists (!!!) and other such "heretics and mad men" rejected by the scientific establishment, will all tell you that they "can do" stuff, and so will all their many paying customers. How do we know they can't do what they say?

Because science gives you the tools to know that you're wrong. If you're a good scientist, you will be wrong _all the time_. That's how science advances: one mistake at a time. But you can't make mistakes if all you ever do is doing stuff with computers, like beating all the benchmarks, because that is a meaningless result judged by its own, self-chosen, measure of success that can never fail; and so can never inform.

Peer review is also not science, but it has been great to catch errors in my papers. Not in conferences, mind. I try to stay away from conferences. Everybody flocks to conferences because of quick turnaround, instant gratification. Journals have the good reviewers who can take their time understanding your work and helping you find where you've gone wrong. "Reject with encouragement to resubmit" is the best review result I ever got.

>> I judge things as "what you can do", not "what can you predict".

The goal of science is not to make predictions, but to understand how the world works, and why. Put that into instrumentalism's pipe and smoke it.

Re: GPT-4 details leaked?

#480

Earlier quoted context omitted.

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science. Machine learning, as it is practiced to…

I'm tired of science as a religion. People treat it as some gospel, like if you check out some criterions you're suddenly "scientific" and instantly get a sense of validity and authority that you shouldn't logically get. I judge things as "what you can do", not "what can you predict". The only demonstration of knowledge and understanding is being able to do something. Not predict. Not "scientific method" and ridiculo…

> In the end of the day, you either manage to do something or you don't.

The hard sciences would like a word.

Post reply on HN