Live data from Hacker News

Training for one trillion parameter model backed by Intel and US govt has begun

techradar.com

151–160 of 267 posts

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#151

Earlier quoted context omitted.

> I bet if you add up that data it's a lot. Let's Fermi estimate that. A 4k video stream is about 50 megabits/second. Let's say that humans have the equivalent of two of those going during waking hours, one for vision and one for everything else. Humans are awake for 18 hours/day, and we'll say a human's training is 'complete' at 25. Multiply that together, and you end up with 1.8e17 bytes, or 180 petabytes of data.…

So much of that data is totally useless though. Most of the data we get is visual and I would argue that visual data is some of the least useful data (with respect to volume). Think about the amount of time we're looking at blank colours with nothing to learn from. Once you've seen one wall, one chair, one table etc, there's not much left to know about those things. An encyclopedia though for example is much less dat…

Also have to disagree here. You see that one object yes, but you see it from 10000000s of different angles and distances. In different settings. In different lighting conditions. With different arrangements. And you see the commonalities and differences. You poke it. prod it. Hit it. Break it. You listen to it.

This is the basis for 'common sense' and I’m pretty sure everything else needs that as the foundation.

Go watch a child learning and you'll see a hell of a lot of this going on. They want to play with the same thing over and over and over again. They want to watch the same movie over and over and over again. Or the same song over and over and over and over and over again.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#152

Earlier quoted context omitted.

Yeah I'm not totally convinced humans don't have a tremendous amount of training data - interacting with the world for years with constant input from all our senses and parental corrections. I bet if you add up that data it's a lot. But once we are partially trained, training more requires a lot less.

> I bet if you add up that data it's a lot. Let's Fermi estimate that. A 4k video stream is about 50 megabits/second. Let's say that humans have the equivalent of two of those going during waking hours, one for vision and one for everything else. Humans are awake for 18 hours/day, and we'll say a human's training is 'complete' at 25. Multiply that together, and you end up with 1.8e17 bytes, or 180 petabytes of data.…

This is completely and fundamentally an incorrect approach from start to finish. The human body - and the human mind - do have electrical and logic components but they are absolutely not digital logic. We do not see in “pixels”. The human mind is an analog process. Analog computing is insanely powerful and exponentially more efficient (time, energy, bandwidth) than digital computing but it is ridiculously hard to pack or compose, difficult to build with, limited in the ability to perform imperative logic with, etc.

You cannot compare what the human eyes/brain/body does with analog facilities to its digital “equivalent” and then “math it” from there.

Also why trying to replicate the human brain with digital logic (current AI approach) is so insanely expensive.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#153

A lot of this was new to me, but it looks like Intel hopes to use this to demonstrate the linear scaling capacity of their Aurora nodes. Argonne installs final components of Aurora supercomputer (22 June 2023): https://www.anl.gov/article/argonne-installs-final-component... Aurora Supercomputer Blade Installation Complete (22 June 2023): https://www.intel.com/content/www/us/en/newsroom/news/aurora... Intel® Data Cent…

yes, I can't imagine the architecture of a supercomputer is the right one for LLM training. But maybe?

If not, spending years to design and build a system for weather and nuke simulations and ending up doing something that's totally not made for these systems is kind of a mind bender.

I can imagine the conversations that led to this: "we need a government owned LLM" "okay what do we need" "lots of compute power" "well we have this new supercomputer just coming online" "not the right kind of compute" "come on, it's top-500!"

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#154
post #117

Earlier quoted context omitted.

The opposite for me. Corporation have naked greed as their driving motivation, but that usually doesn't involve killing off all of their customers. That would be quite unprofitable. People elected to government often seem to seek power for powers sake and I have less faith that they'll not harm us.

Such as when energy companies buried research about climate change...?

That seems to pale in comparison to what governments have done.

Doing more to stem the growth in CO2 emissions would have reduced the magnitude of the change in the climate, and with it, the harmful effects of that change, but it would also have reduced the benefits that oil and gas have conferred upon the world, in raising incomes and reducing poverty worldwide.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#155
post #114

Earlier quoted context omitted.

Uhm the east indian company has an estimated death toll in the dozens of million people? [1] [1] https://en.m.wikipedia.org/wiki/Timeline_of_major_famines_in... The company was only nationalised in 1858 and until then effectively colonised india until then

> [The Queen] granted her charter to their corporation named Governor and Company of Merchants of London trading into the East Indies.[15] For a period of fifteen years, the charter awarded the company a monopoly[26] on English trade with all countries east of the Cape of Good Hope and west of the Straits of Magellan.[27] Any traders there without a licence from the company were liable to forfeiture of their ships an…

Monopoly is the natural state. Government is the only reason we have any alternatives.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#156
post #117

Earlier quoted context omitted.

The opposite for me. Corporation have naked greed as their driving motivation, but that usually doesn't involve killing off all of their customers. That would be quite unprofitable. People elected to government often seem to seek power for powers sake and I have less faith that they'll not harm us.

Such as when energy companies buried research about climate change...?

Then again it was the Greens that blocked nuclear power.

In Germany, their actions actually caused increase in the use of coal and increased CO2 output.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#157

Earlier quoted context omitted.

Ideally no one person or entity controls such a thing. But, would I rather have a Government, or a corporation control AGI? If I had to pick one of two evils, the Government would be the lesser of the two.

The opposite for me. Corporation have naked greed as their driving motivation, but that usually doesn't involve killing off all of their customers. That would be quite unprofitable. People elected to government often seem to seek power for powers sake and I have less faith that they'll not harm us.

I sure see a lot of power for power’s sake inside of companies.

OpenAI’s stated mission is AGI that can replace half the population in “economically valuable” work.

I get that with some creativity you can see that as a net benefit for humanity, but across at least a generation that’s going to be a rough & destabilizing transition.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#158

I'm sure the government's mission is also to develop an AGI that benefits us all.

Ideally no one person or entity controls such a thing. But, would I rather have a Government, or a corporation control AGI? If I had to pick one of two evils, the Government would be the lesser of the two.

Yeah but AGI would control both in the end.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#159

Earlier quoted context omitted.

Ideally no one person or entity controls such a thing. But, would I rather have a Government, or a corporation control AGI? If I had to pick one of two evils, the Government would be the lesser of the two.

Yeah but AGI would control both in the end.

Looking at our two leading presidential candidates could that be an improvement?

Maybe AGI takes over after it emerges as a solution to our inability to select competent candidates for political office.

Re: Training for one trillion parameter model backed by Intel and US govt has begun

#160
post #5

Isn't GPT-4 already over 1T parameters? And GPT-5 should be even "an order of magnitude" bigger than GPT-4...

Was it ever confirmed whether GPT-4 is a Mixture of Experts or not?

I just asked GPT-4, and it denies being a MoE model.

User:

Are you an MoE model?

ChatGPT:

No, I am not based on a Mixture of Experts (MoE) model. My underlying architecture is based on the GPT (Generative Pre-trained Transformer) framework, specifically the GPT-4 version. This architecture is a large-scale transformer-based neural network, but it does not use the MoE approach.

In a GPT model like mine, the entire model works in a unified manner on each task, without dividing it among specialized sub-models or experts. The model is trained on a diverse range of data and tasks, enabling it to generate responses across a wide array of topics and questions. While this approach lacks the specialization that MoE models offer, it provides broad, general-purpose language understanding and generation capabilities.

Post reply on HN