Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

221–230 of 648 posts

Re: GPT-4 details leaked?

#221

I wonder what the legal implications of them using SciHub and Libgen would be if that's true. I'd imagine OpenAI is big enough to make deals with publishers.

probably just easier to use drm-free copies of books

Re: GPT-4 details leaked?

#222

Earlier quoted context omitted.

4-bit quantization removes a lot of the model's sophistication, and 60B parameters is still smaller than what GPT4 is using.

The point is that it's infinitely better in not being there "just to take your jobs and make a few VCs richer". Nobody even claimed it's more performant. It's like the difference getting nothing, but keeping your land, and getting glass pearls, losing your land. You have to completely ignore the meat of the argument to even pretend there is a contest. And this is without considering what happened if we stopped feedin…

There is no way I am going to spin up my own worse LLM so a few people will make less money. Even if it was 1-5% better. It's just not worth the time.

Re: GPT-4 details leaked?

#223

Earlier quoted context omitted.

From the post: [Rumors that start to become lawsuits] Some speculations are: - LibGen (4M+ books) - Sci-Hub (80M+ papers) - All of GitHub This is the most funny, but in the end sad aspect. If ChatGPT was indeed trained on pirated content and is able to be(come) such a powerful tool, then the copyright laws should have been abolished yesterday. If ChatGPT was not trained on all these resources out there, then think ho…

That is a very dangerous way to approach copyright laws. They are definitely abused by corporations like Disney, infamously so, but abolishing them is absolutely not the answer. Art makers are already struggling en masse, taking away their ability to earn money off their work isn't an answer, especially if it's just to train a predictive text generator.

The laws themselves are dangerous and encourages the myth of ideas being akin to somebody's personal property.

Re: GPT-4 details leaked?

#224

Earlier quoted context omitted.

LLaMA 30B or 60B can be very impressive when correctly prompted. Deploying the 60B version is a challenge though and you might need to apply 4-bit quantization with something like https://github.com/PanQiWei/AutoGPTQ or https://github.com/qwopqwop200/GPTQ-for-LLaMa . Then you can improve the inference speed by using https://github.com/turboderp/exllama . If you prefer to use an "instruct" model à la ChatGPT (i.e. tha…

4-bit quantization removes a lot of the model's sophistication, and 60B parameters is still smaller than what GPT4 is using.

> 60B parameters is still smaller than what GPT4 is using

I mean if the article is right, then it's about 3.3% the size of GPT 4 (although it's a sparse model so not all of it is used on every pass).

Meta also didn't train LLaMAs on nearly as much code it seems, so they're much worse for that in general.

Re: GPT-4 details leaked?

#225
post #79

Earlier quoted context omitted.

Unfortunately I've found the current OSS models to be vastly inferior to the OpenAI models. Would love to see someone actually get close to what they can do with GPT-3.5/4, except capable of running on commodity GPUs. What's the most impressive open model so far?

Have you tried falcon 40b instruct? Also take into account that chatgpt likely has some preprompt and by talking to falcon or other OS models it's all in your hands. Furthermore, Not many people discuss the significance of proper output sampling. I myself used to just test open source models with the greedy decoding only. Who knows if they wouldn't even beat (not at all)OpenAI with some clever output sampling scheme.

Idk, has anyone tried falcon yet? The support for running it remains nonexistent except for one fork of llama.cpp that isn't integrated into anything. This trend of every new model breaking compatibility really needs to stop.

Re: GPT-4 details leaked?

#226
post #182

Earlier quoted context omitted.

When choosing titles for my own submissions, yeah, the accurate title that HN says they desire gets no votes whatsoever. Any clickbait on here, people bring upon themselves (and this isn't even a clickbait-level title)

I don't want accurate titles because they'll make me vote for it. I want accurate titles because it helps me determine if I'll read it BEFORE clicking it. The whole point of accurate titles is that you'll get less votes on uninteresting content.

But if the title requires you to click, and then you find out it's uninteresting, why'd you upvote at that point? It shouldn't get your vote at all then, having wasted your time

Re: GPT-4 details leaked?

#227
post #209
post #187

Earlier quoted context omitted.

wont that nearly kill your ssd if you do it for extended periods of time?

Most of the ram is for storing the model once it is loaded it is read only so will not harm an SSD.

It only reads from memory,not swap directly. If it needs to read something from swap, it'll write out something from memory to swap, then read the swap into memory. Reading 1gb of swap, will essentially write 1gb to the ssd too. (rough numbers)

Correct me if I misunderstand swap?

Re: GPT-4 details leaked?

#228

"Open" AI, a charity to benefit us all by pushing and publishing the frontier of scientific knowledge. Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. https://github.com/ggerganov/llama.cpp https://github.com/openlm-research/open_llama https://huggingface.co/TheBlok…

> try to pressure the government into making it illegal for you to compete with us. I mean the guy who created GPT-4 literally demanded a ban of any system more powerful than GPT-4.

I don’t understand this. Won’t that hurt their progress on GPT-5?

Re: GPT-4 details leaked?

#229

"Open" AI, a charity to benefit us all by pushing and publishing the frontier of scientific knowledge. Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. https://github.com/ggerganov/llama.cpp https://github.com/openlm-research/open_llama https://huggingface.co/TheBlok…

>> We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us.

Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science.

Machine learning, as it is practiced today, is not science. There is no scientific theory behind it and there is no scientific method applied. There are no scientific questions asked, or attempted to be answered. There is no new knowledge produced other than how to tune systems to beat benchmarks. The standard machine learning paper is a bunch of text and arcane-looking formulae around a glorified leaderboard: a little table with competing systems on one side and arbitrarily chosen benchmark datasets on the other side; and all our results in bold so everyone knows we're winning. That's as much doing science as is racing cool-looking sports cars.

Re: GPT-4 details leaked?

#230

>If their cost in the cloud was about $1 per A100 hour, the training costs for this run alone would be about $63 million. If someone legitimate put together a crowd funding effort, I would donate a non-insignificant amount to train an open model. Has it been tried before?

Not yet, heard tale of several people having the same idea to train an open model though through either crowdfunding or some wizardry with crowdsourcing GPUs. $65 million sounds pretty high though.

Considering an effort to buy a copy of the constitution raised almost $47M, I wouldn't be so sure. [^1]

Worth noting, though, that it isn't just the computing budget that's missing here - it is also (and perhaps even more importantly) the high quality data to actually train the model.

[^1]: https://en.wikipedia.org/wiki/ConstitutionDAO

Post reply on HN