Live data from Hacker News

Training for one trillion parameter model backed by Intel and US govt has begun

techradar.com

241–250 of 267 posts

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#241

Earlier quoted context omitted.

Monopoly is the natural state. Government is the only reason we have any alternatives.

Monopoly is the byproduct of allowing centralized power, not the natural state. I'm not actually sure how we could narrow down the natural state of humans at this point, but I strongly suspect it wouldn't be based on an assumption that people are willing to give up a growing list of individual freedoms in the name of fear.

Stopping the accumulation of power requires aggressive sacrifice from the less powerful.

This isn’t a feature of humans, but basic system dynamics / economics / etc.

A group gets more power and leverage that power to gain more power. Inevitable. Coordinated action is the only way to prevent it. And coordinated action is government.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#242

Earlier quoted context omitted.

LLMs don’t have insight, these outputs cannot be assumed to be accurate. Due to current hardware limitations it is not feasible to have a 1T parameter model without MoE.

I wonder if it has some canned human-written responses when asking specific questions about itself. This would be pretty clever to silently implement, that will definitely help convince people that it's approaching "AGI". It's possible that it's just hallucinating here too, I don't have any proof that the responses are canned, but they appear that way to me.

Yes, it has an invisible "shadow prompt" that gets sent to it when you start a session. It looks something like this:

https://www.reddit.com/r/ChatGPT/comments/zo9of4/comment/j0n...

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#243
Is anything known about what extent if any non-public domain books are used for LLM’s?

One example is the Google books project made digital quite a few texts, but I’ve never heard if Google considers these fair game to train on for Bard.

Most of the copyright discussions I’ve seen have been around images and code but not much about books.

Seems to become more relevant as things scale up as indicated by this article.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#244

Earlier quoted context omitted.

Curious, why do people gravitate to the term "hallucination" (as if it's a person not a tool) instead of just plainly stating that it's wrong?

If I ask it what color an orange it and it says blue, that would be wrong. If you ask it a question and it makes up a completely fabricated story, like for example the case files in that recent legal case [1], then saying it was “wrong” doesn’t really seem to capture it. Calling it a hallucination is a great analogy, because the model made up a plausible sounding, but completely fabricated story. It saw things that w…

Fair enough and thanks. I felt hallucination was too forgiving a term but I can see how others would rank them the other way around and suppose it works.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#245

Earlier quoted context omitted.

Holy shit, you are right. They probably have 10-100x the data used to train gpt-4. Decades of every text message, phone call transcript, and so on. I can’t believe I haven’t seen anyone mention that yet. People keep saying we don’t have enough data. I think there is a lot more data than we realize, even ignoring things like NSA.

Apparently there are roughly 2 trillion text messages sent per year in the US [1]. I did a sanity check, that’s like 40 or so a day per person, so sounds reasonable. I couldn’t find the average message length, but I would guess it’s fairly short (with a fat tail of longer messages). To make the math easy, let’s say the average length is ~10 tokens. I’d be surprised if that isn’t correct within a factor of 2 or so. So…

Interestingly, the reason Google initially created it's Google Voice service back in the day was to gather voicemail audio to train its speech to text engines.

It's mind-blowing to me that with all of Google's data, Google isn't the far and away leader in this new space. I have to believe they're paralyzed by the fear of legal repercussions.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#246

Earlier quoted context omitted.

Monopoly is the byproduct of allowing centralized power, not the natural state. I'm not actually sure how we could narrow down the natural state of humans at this point, but I strongly suspect it wouldn't be based on an assumption that people are willing to give up a growing list of individual freedoms in the name of fear.

Stopping the accumulation of power requires aggressive sacrifice from the less powerful. This isn’t a feature of humans, but basic system dynamics / economics / etc. A group gets more power and leverage that power to gain more power. Inevitable. Coordinated action is the only way to prevent it. And coordinated action is government.

It sounds like you're describing coordinated action as both the cause of and solution to the same problem.

Stopping this accumulation of power takes very little, its undoing the accumulation of power that is costly.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#247

Earlier quoted context omitted.

I just realized that the NSA has probably been able to train GPT-4 equivalents on _all the data_ for a while now. We'll probably never learn about it but that's maybe scarier than just the Snowden collection story because LLMs are so good at retrieval.

Holy shit, you are right. They probably have 10-100x the data used to train gpt-4. Decades of every text message, phone call transcript, and so on. I can’t believe I haven’t seen anyone mention that yet. People keep saying we don’t have enough data. I think there is a lot more data than we realize, even ignoring things like NSA.

[deleted]

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#248

Earlier quoted context omitted.

Was it ever confirmed whether GPT-4 is a Mixture of Experts or not?

I just asked GPT-4, and it denies being a MoE model. User: Are you an MoE model? ChatGPT: No, I am not based on a Mixture of Experts (MoE) model. My underlying architecture is based on the GPT (Generative Pre-trained Transformer) framework, specifically the GPT-4 version. This architecture is a large-scale transformer-based neural network, but it does not use the MoE approach. In a GPT model like mine, the entire mod…

Can you explain how your own mind works?

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#250
post #167

Earlier quoted context omitted.

> I bet if you add up that data it's a lot. Let's Fermi estimate that. A 4k video stream is about 50 megabits/second. Let's say that humans have the equivalent of two of those going during waking hours, one for vision and one for everything else. Humans are awake for 18 hours/day, and we'll say a human's training is 'complete' at 25. Multiply that together, and you end up with 1.8e17 bytes, or 180 petabytes of data.…

Nit: The bandwidth of your optic nerve is only about 10 kilobits per second. You think you're seeing in 4K but most of it is synthetic.

That cannot be right.

10 kilobits per second would only get you some awful telephone quality audio. Even super low res video would be orders of magnitude more than that.

Post reply on HN