Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

361–370 of 648 posts

Re: GPT-4 details leaked?

#361
post #211

Earlier quoted context omitted.

The T in ChatGPT stands for Transformers. The similarity between the OG Transformer from 2017 and GPT3 (and other modern LLMs) is pretty big

The data, size and training process are what's different.

The point is LLMs are built with Transformer architecture. They're all transformers, attention is an integral part of building worthwhile, contextual answers.

Re: GPT-4 details leaked?

#362

Earlier quoted context omitted.

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science. Machine learning, as it is practiced to…

It's more a science than pure math or CS, since they are producing falsifiable models.

Exactly. I think the word "science" changed its meaning a lot these years.

For some reason people tend to consider a field that is more "formal" (like pure math, some CS concepts like lambda calculus) more science, even though historically formal systems came very late, and practically very few systems can be described that way.

I really wonder whether people who regurgitate "machine learning isn't science" think theory of evolution is science or not.

Re: GPT-4 details leaked?

#363
post #355

"Open" AI, a charity to benefit us all by pushing and publishing the frontier of scientific knowledge. Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. https://github.com/ggerganov/llama.cpp https://github.com/openlm-research/open_llama https://huggingface.co/TheBlok…

Wow, the top comment is neither relevant to the post, nor friendly or interesting. Activism, even with false premises. Many of us tried, and those with a little sense left know that running your local LLM on a non-GPU is not really useful. Besides, what does your post add to the discussion, and why is it the top posting? Create your local LLM, use it, tell other people about how you did it exactly, and be happy. But…

Well in 1929-1930 The stock market crashed and regulating others in a market became needed. Now it’s who can regulate the market best in their interests, like cooperations dominate with bailouts and subsidies. You’re picking a fight with a small actor while OpenAI is spending millions doing the same, fighting for regulating to keep their status quo

Re: GPT-4 details leaked?

#364
post #139

Earlier quoted context omitted.

In today's world, "Science equals Capitalism". Or at least Science is allowed to progress and get funded only as long as it serves the interest of Capitalism.

That's engineering, not science.

The National Science Foundation is the architect behind the whole STEM Pipeline. It’s one in the same

Re: GPT-4 details leaked?

#365

Earlier quoted context omitted.

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science. Machine learning, as it is practiced to…

{Hypothesis, test, loop} is the scientific method, and I can guarantee it is being used when fine tuning an LLM.

That's a common interpretation of what science is, but it largely ends up being driven by confirmation bias. Because how do you know your hypothesis and test are even really connected? Or what you're seeing is a cause and not a correlation? This is why you need arguably the two most important factors in "real" science: predictability and falsifiability.

Predictability means that if your hypothesis is correct then you'd be able to formulate other improbable (and ideally currently untestable) predictions from it. And falsifiability means that if these predictions fail to occur then your initial hypothesis was also almost certainly wrong. So for instance Newton's hypothesis was that gravity was driven by a mathematical relationship between the mass/distance of two bodies. It was good science because it lead to the shocking ability to be able to dramatically simplify orbital dynamics, and create a complete predictive system of these bodies. It was even used to mathematically discover a completely unknown planet - Neptune.

Incidentally, his theory would also be able to be shown to be false if any of these unexpected predictions ended up being false. And that's actually exactly what happened. Observation of Mercury's orbit about the Sun showed it was off by ~1/3600th of one degree per century, relative to what was expected. And it's from there that people knew there was a mistake, which would only be explained later by Einstein hundreds of years later, who hypothesized a system with far more absurd predictions... and so the story continues.

Re: GPT-4 details leaked?

#366

Earlier quoted context omitted.

It's like asking whether a carpenter is a scientist because they developed a cabinet...

I hate to break it to you but that is part of science. Perhaps the major part of it too. Hypothesis: I can build a cabinet with these materials which will bear some load range. Experiment: I built it and it obviously works. Now change cabinet to particle accelerator that a giant team of other theorists and engineers designed. Am I not doing science by participating in building it? So experimental scientists arent sci…

To me, you're describing the differences between a cook and a chef. Just because you've built something with well known methods doesn't make you a scientist. You didn't come up with anything new. In fact, we're starting to sound a lot like Apple. Apple is (in)famous for taking ideas that someone else did all of the hard work of developing and proving to work, and then take various ideas like that to combine into an actual useful working something. (They also do a lot of pure deep research as well, so don't think I'm going too far with the analogy.)

Re: GPT-4 details leaked?

#367
post #9

Previously posted about here: https://news.ycombinator.com/item?id=36671588 and here: https://news.ycombinator.com/item?id=36674905 With the original source being: https://www.semianalysis.com/p/gpt-4-architecture-infrastruc... The twitter guy seems to just be paraphrasing the actual blog post? That's presumably why the tweets are now deleted. --- The fact that they're using MoE was news to me and very interesting. I…

The previous posts are to a twitter thread that's been taken down, and the preview of a post that requires a $1000 subscription. This post however is freely available (for now at least).

And the tweeter of the twitter thread paid the $1000, copied the useful info to twitter, and then did a credit card chargeback.

Re: GPT-4 details leaked?

#368

"Open" AI, a charity to benefit us all by pushing and publishing the frontier of scientific knowledge. Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. https://github.com/ggerganov/llama.cpp https://github.com/openlm-research/open_llama https://huggingface.co/TheBlok…

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science. Machine learning, as it is practiced to…

Science is knowing how the model scales when you throw this much processing power at it. They won't even tell us how much processing power that is.

Re: GPT-4 details leaked?

#369
post #167

Earlier quoted context omitted.

Yes, they are doing the improving, but then you need loads of money to do the learning no university can afford. So now big tech is hiring promising university researchers for good money to scale up their research. This could be solved by massive decentralization where millions of users provide compute with their gpus and i think it will be at some point, cause i believe foss is more powerful than this openai bs. The…

> There are people working on this, but afaik the techniques aren't quite there. You need a different kind of model with much more parallelization then what is currently used. What if crypto is switching from mindless hashing as proof-of-work to training AI models as proof-of-work? That would mean suddenly big computing resources are available.

“What if running part of the paperclip maximiser that destroys humankind was profitable?” is both a bit of a nightmare sci-fi dystopia, and a reasonable description of companies in the modern economy.

Re: GPT-4 details leaked?

#370

Earlier quoted context omitted.

LLaMA 30B or 60B can be very impressive when correctly prompted. Deploying the 60B version is a challenge though and you might need to apply 4-bit quantization with something like https://github.com/PanQiWei/AutoGPTQ or https://github.com/qwopqwop200/GPTQ-for-LLaMa . Then you can improve the inference speed by using https://github.com/turboderp/exllama . If you prefer to use an "instruct" model à la ChatGPT (i.e. tha…

Just a reminder that LLaMA is not open—in order to use it legally you have to agree to Meta's terms, which currently means research use only. The versions circulating on torrents are essential pirated, and while I don't have an ethical problem with that at all you can't use it safely in a business. The open replacements for LLaMA have yet to reach 30B, let alone 65B.

If anyone has a copyright claim to an LLM, the creators of the input data have more of a copyright claim than the company that trained it. There's a good chance they are not copyrightable at all. I'd bet there's a lot of people willing to take on that risk.

However, they might still fall under trade secret law.

Post reply on HN