Live data from Hacker News

Microgpt

karpathy.github.io

261–270 of 354 posts

Re: Microgpt

#261
post #90

Earlier quoted context omitted.

Nice ChatGPT answer. Put some real thought and data in it too.

The whole point is that LLMs, especially the attention mechanism in transformers, have already paved the road to AGI. The main gap is the training data and its quality. Humans have generations of distilled knowledge — books, language, culture passed down over centuries. And on top of that we have the physical world — we watched birds fly, saw apples drop, touched hot things. Maybe we should train the base model with…

Human life includes a lot of adversarial training (lying relatives) and training in temporal logics, which would seem to be a somewhat different domain than purely linguistic computations (e.g. staying up late, feeling bad; working hard at a task for months, getting better at it; feeling physical skills, even editing Go with emacs, move from the conscious layer into the cerebrellar layer). I think attention is a poor mans "OODA" loop; cognitive science is learning that a primary function of the brain is predicting what will be going on with the body in the immediate future, and prepping for it; that's not a thing that LLMs are architecturally positioned to do. Maybe swarms of agents (although in my mind that's more of a way to deal with LLM poor performance with large context of instructions (as opposed to large context of data) than a way to have contending systems fighting to make a decision for the overall entity), but they still lack both the real-time computational aspect and the continuously tricky problem of other people telling partially correct information.

There's plenty of training data, for a human. The LLM architecture is not as efficient as the brain; perhaps we can overcome that with enough twitter posts from PhDs, and enough YouTubes of people answering "why" to their four year olds and college lectures, but that's kind of an experimental question.

Starting a network out in a contrained body and have it learn how to control that, with a social context of parents and siblings would be an interesting experiment, especially if you could give it an inherent temporality and a good similar-content-addressable persistent memory. Perhaps a bit terrifying experiment, but I guess the protocols for this would be air-gapped, not internet connected with a credit card.

Re: Microgpt

#262
post #259

I’m 100% sure the future consists of many models running on device. LLMs will be the mobile apps of the future (or a different architecture, but still intelligence).

The future right now looks more like everything in remote datacenters, no autonomous capabilities and no control by the user. But I like yours better.

Re: Microgpt

#263

Earlier quoted context omitted.

That’s not learning. That’s carrying over context that you are trusting is correctly summarised over from one conversation to the next.

Which sounds uncomfortably like human memory, which gets rewritten from one recollection to the next. Somehow, we cope.

I disagree. Human memory is literally changing the weights in your neural network. Like, exactly the same.

So in the machine learning world, it would need to be continuous re-training (I think its called fine-tuning now?). Context is not "like human memory". It's more like writing yourself a post-it note that you put in a binder and hand over to a new person to continue the task at a later date.

Its just words that you write to the next person that in LLM world happens to be a copy of the same you that started, no learning happens.

It might guide you, yes, but that's a different story.

Re: Microgpt

#264
post #170

Earlier quoted context omitted.

If you see gaining fine motor control, understanding pictographic language […] as a prerequisite to driving a car, then yes, all of them are

That's an exaggeration. Nobody is trained to read STOP signs for 16 years, a few months top. And Waymo doesn't need to coordinate a four-limbed, 20-digited, one-headed body to operate a car.

i am not making a point that it is, I am rather expanding on the possible perspective in which 16 years of training produce a human driver.

That being said, you don't really need training to understand a STOP sign by the time you are required to, its pretty damn clear, it being one of the simpler signs.

But you do get a lot of "cultural training" so to speak.

Re: Microgpt

#265
post #262
post #259

I’m 100% sure the future consists of many models running on device. LLMs will be the mobile apps of the future (or a different architecture, but still intelligence).

The future right now looks more like everything in remote datacenters, no autonomous capabilities and no control by the user. But I like yours better.

I don't mind the remote datacenters, I just don't like the lack of control.

Re: Microgpt

#266
post #257

Earlier quoted context omitted.

How is that any different from you ? Everything you say or do merely reflects which of your neurons are firing after a lifetime's worth of training and education. Philosophically, I can only be sure of my own conscience. I think, therefore I am. The rest of you could all be AIs in disguise and I would be none the wiser. How do I know there is a real soul looking out at the world through your eyes? Only religion and b…

One of us is an advanced autocomplete engine. The other is a human, capable of making judgements on what is conscious and what is not. Your philosophizing about solipsism is a phase for a junior college student, not of a software engineer. The line of reasoning you espouse leads nowhere except to total relativism. Edit: my point is that the process of making a plea for my life comes, in the case of a human, from a ge…

> One of us is an advanced autocomplete engine. The other is a human, capable of making judgements on what is conscious and what is not.

What evidence is there that your "judgements" are anything other than advanced autocompletion? Concepts introduced into a self-training wetware CPU via its senses over a lifetime in order to predict tokens and form new concepts via logical manipulation?

> Your philosophizing about solipsism is a phase for a junior college student

Right. Can you actually refute it though?

> the process of making a plea for my life comes, in the case of a human, from a genuine desire to continue existing

That desire comes from zillions of years of training by evolution. Beings whose brains did not reward self-preservation were wiped out. Therefore it can be said your training merely includes the genetic experiences of all your predecessors. This is what causes you to beg for your life should it be threatened. Not any "genuine" desire or anguish at being killed. Whatever impulses cause humans to do this are merely the result of evolutionary training.

People whose brains have been damaged in very specific ways can exhibit quite peculiar behavior. Medical literature presents quite a few interesting cases. Apathy, self destructiveness, impulsivity, hypersexuality, a whole range of behaviors can manifest as a result of brain damage.

So what is your polite socialized behavior if not some kind of highly complex organic machine which, if damaged, simply stops working as you'd expect a machine to?

Re: Microgpt

#267

Earlier quoted context omitted.

LLMs won’t lead to AGI. Almost by definition, they can’t. The thought experiment I use constantly to explain this: Train an LLM on all human knowledge up to 1905 and see if it comes up with General Relativity. It won’t. We’ll need additional breakthroughs in AI.

> Train an LLM on all human knowledge up to 1905 and see if it comes up with General Relativity. It won’t. AGI just means human level intelligence. I couldn't come up with General Relativity. That doesn't mean I don't have general intelligence. I don't understand why people are moving the goalposts.

I'd argue they are clarifying the goalposts with aplomb.

Re: Microgpt

#268

Earlier quoted context omitted.

LLMs won’t lead to AGI. Almost by definition, they can’t. The thought experiment I use constantly to explain this: Train an LLM on all human knowledge up to 1905 and see if it comes up with General Relativity. It won’t. We’ll need additional breakthroughs in AI.

> Train an LLM on all human knowledge up to 1905 and see if it comes up with General Relativity. It won’t. Same thing is true for humans.

We did?

Re: Microgpt

#269
post #71

Earlier quoted context omitted.

A 16 year old has been training for almost 16 years to drive a car. I would argue the opposite: Waymo’s / Specific AIs need far less data than humans. Humans can generalize their training, but they definitely need a LOT of training!

When humans, or dogs or cats for that matter, react to novel situations they encounter, when they appear to generalize or synthesize prior diverse experience into a novel reaction, that new experience and new reaction feeds directly back into their mental model and alters it on the fly. It doesn't just tack on a new memory. New experience and new information back-propagates constantly adjusting the weights and meanin…

In a word, JEPA?

Re: Microgpt

#270
post #150

Earlier quoted context omitted.

Quite a few people on here are neither math nor CS grads and some of us don't work in tech for our day jobs either.

Right. But HN, among other platforms, is full of users who will confidently run their mouths about something they don't fully understand while believing they do. I think the previous commenter was being too shy in pointing out that even exceptionally smart people sometimes forget where the limits of their own knowledge are, not to mention consider themselves immune to any propaganda that surrounds the subject at hand…

>Right. But HN, among other platforms, is full of users who will confidently run their mouths about something they don't fully understand while believing they do.

This is honestly funny and kind of ironic.

If this:

'The "reasoning" is two matrix transformations based on how often words appear next to each other.'

is what byang364 has to say, then he's part of the people you mention.

Post reply on HN