Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

341–350 of 648 posts

Re: GPT-4 details leaked?

#341
post #315

Earlier quoted context omitted.

MS not owning OpenAI outright is a technicality. "It would suck to be an MS shareholder" due to things done to subsidiaries of MS constitutes dereliction of duty by MS officers. MS is not Silicon Valley. They're based in Seattle and predate 95% of SV. Besides this is just one of my points. Even if everything you claimed were to be true, OpenAI still is, as I pointed out, part of the American Corporate establishment a…

> OpenAI still is, as I pointed out, part of the American Corporate establishment and despite their PR posturing, they won't provide us with technology to disrupt said establishment They're not acting like they're part of the American corporate establishment, and I don't understand why you think they are. Rather than quote my other comment comparing against 3M, I'll link to it: https://news.ycombinator.com/item?id=36…

[flagged]

Re: GPT-4 details leaked?

#342
post #322
post #317

Earlier quoted context omitted.

> The written testimony didn't say "ban", but it did stress fairly complex regulation You feel that quotation justifies the adjective "complex"?

I didn't mean "complex" in some sort of flippant, derogatory way. The language is suggesting regulation with active government involvement and oversight (licensing, for example). In the range of regulations in the software world, that would be on the far right end. ITAR, for example, requires licensing. People regularly describe the process as "complex".

Hmm.

While I don't doubt that ITAR is especially complex, following legislation is pretty common as a basic requirement in software development — say GDPR, or tax rules, or COPPA, etc. — it really doesn't seem to me that this is pushing all that hard for anything specific and detailed enough to even justify the claim that they're calling for a specific level of complexity of legislation.

Re: GPT-4 details leaked?

#343
post #342
post #322

Earlier quoted context omitted.

I didn't mean "complex" in some sort of flippant, derogatory way. The language is suggesting regulation with active government involvement and oversight (licensing, for example). In the range of regulations in the software world, that would be on the far right end. ITAR, for example, requires licensing. People regularly describe the process as "complex".

Hmm. While I don't doubt that ITAR is especially complex, following legislation is pretty common as a basic requirement in software development — say GDPR, or tax rules, or COPPA, etc. — it really doesn't seem to me that this is pushing all that hard for anything specific and detailed enough to even justify the claim that they're calling for a specific level of complexity of legislation.

Yeah, I'm not trying to argue that what's being proposed is bad, just trying to clarify what he said. The licensing part does make it more like ITAR and less like GDPR, HIPPA, etc. Those sorts of laws don't have active government activity to certify and license beforehand, though there are 3rd party orgs that will do that sort of thing if you want.

Re: GPT-4 details leaked?

#344

Earlier quoted context omitted.

The difference is that you can reason about why it didn't work. With deep learning models not so. They are too big to reason about in the same sense as your BST algorithm. Hence, you need a scientific approach to construct them. I.e., with lots of experimenting, hypotheses, etc.

That seems like it pushes it further from science, no? The point of a well-crafted hypothesis is that if it doesn’t bear out, you know that it’s because one+ of your assumptions was wrong. Your ability to continue your scientific inquiry is pretty much == your ability to then identify which assumption was wrong.

You don't need to do an scientific experiment to tell why your BST doesn't work. "Computer Science" is a misnomer because most of its contents and methodologies are from mathematics (which is used in science but is not a science in itself).

CS uses mathematical proofs. You don't need a computer to execute your code to tell why the BST doesn't work. You can introspect your code and figure out why it works or does not work. If it's correct, CS methodology says you can "prove" that it works (without executing it).

Working with large AI models is like working with an artificial brain. It's as scientific as neuroscience in this sense. You make some hypotheses, tweak some hyperparameters, and get a result, which may or may not invalidate your hypotheses. Nobody knows why. Science is not necessarily about knowing the fundamental "whys" (amateurs think humanity has figured all the "whys" out, but that's a lie). It's about establishing some useful model of how things work.

But it's definitely possible to know why your BST does not work, even without a computer, without empirical testing. That's why CS is not a science.

Re: GPT-4 details leaked?

#345
post #135

If this is true, then: 1. Training took 21 yottaflops. When was the last time you saw the yotta- prefix for anything? 2. The training cost of GPT-4 is now only 1/3 of what it was about a year ago. It is absolutely staggering how quickly the price of training an LLM is dropping, which is great news for open source. The google memo was right about the lack of a moat.

>> The training cost of GPT-4 is now only 1/3 of what it was about a year ago. It is absolutely staggering how quickly the price of training an LLM is dropping, which is great news for open source. The google memo was right about the lack of a moat. That really doesn't change anything at all. The more training large models gets cheaper, the more large corporations are able to train larger models than everyone else. S…

Training data quality and quantity is the bottleneck.

"Chinchilla showed that we need to be using 11× more data during training than that used for GPT-3 and similar models. This means that we need to source, clean, and filter to around 33TB of text data for a 1T-parameter model." https://lifearchitect.ai/chinchilla/

GPT4 has been trained on images exactly for this reason (it might not have been worth it separately from multi-modality, but together these two advantages seem decisive).

Re: GPT-4 details leaked?

#346

Earlier quoted context omitted.

The real moat is an abundance of high quality data.

Well open AI raised eye brows by crawling the internet and using everyone's data to make a commercial product One day some new startup will train on all of libgen and torrent networks, but it will be very hard to prove. You'll keep getting these gaps up in questionable morality and legality, and even openai will complain about playing fair

ThePile already contains some content from a torrent, and there's as lawsuit alleging that Meta has committed copyright infringement by using it.

https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o...

Re: GPT-4 details leaked?

#348

Earlier quoted context omitted.

I'm not sure why you are not understanding that not getting residuals is not the same as not getting a salary or a per writing contribution payment. Or that production companies are not paying these salaries (and the rest of the production budget) out of the kindness of their hearts, but out of revenues accrued from networks and streaming services having to pay to screen their shows. Or that writers will not be paid…

If it's genuinely news to you that exploitation is a normal day-to-day part of any industry, perhaps these pointless replies should end here. Yes, writers have been working for free [1]. [1] "In October 2015, Wil Wheaton created a stir when he declared that he had turned down an offer to write for the Huffington Post. He refused, according to him, because they had declined to pay for his work, in keeping with their p…

Strangely, it is not news to me that exploitation is a normal day-to-day part of any industry .

(This is why I have not made any statements to that effect, never mind done anything as ludicrous as post articles about unionised screenwriters seeking to negotiate a higher pay rate as evidence that "most" of them earned "nothing")

And this is also why I am not advocating a copyright-free world in which HuffPo has the right to sell ads around everything Wil Wheaton ever wrote without paying him a penny or even seeking his permission. Guess what writers whose work isn't copyrightable will be paid in? That's right, "exposure", and not even exposure with much prospect of paid compensation if their work takes off.

Re: GPT-4 details leaked?

#349

Earlier quoted context omitted.

...but in a technical context, not in a scientific one. One thing is identifying (age of copper; age of bronze) the best ways of smelting ore to obtain the metal through trial and error, another is to try and understand the nature of materials.

Those steps are both necessary and usually take decades to build understanding.

Next time we get the urge to complain that GPT-n is just applying patterns seen in the training corpus we should remember humans don't do much better. We are just language agents with rich feedback from outside. We can't even write software top-down in one go without running the code.

Re: GPT-4 details leaked?

#350

Earlier quoted context omitted.

Unfortunately I've found the current OSS models to be vastly inferior to the OpenAI models. Would love to see someone actually get close to what they can do with GPT-3.5/4, except capable of running on commodity GPUs. What's the most impressive open model so far?

LLaMA 30B or 60B can be very impressive when correctly prompted. Deploying the 60B version is a challenge though and you might need to apply 4-bit quantization with something like https://github.com/PanQiWei/AutoGPTQ or https://github.com/qwopqwop200/GPTQ-for-LLaMa . Then you can improve the inference speed by using https://github.com/turboderp/exllama . If you prefer to use an "instruct" model à la ChatGPT (i.e. tha…

How do you correctly prompt it? A lot of people are not familiar with how to do this. I think this would improve how many people are using the non-OpenAI models.
Post reply on HN