Live data from Hacker News

Executing programs inside transformers with exponentially faster inference

percepta.ai

131–139 of 139 posts

Re: Executing programs inside transformers with exponentially faster inference

#131

Earlier quoted context omitted.

Full disclosure: all my published work is on symbolic machine learning (a.k.a. Inductive Logic Programming) :O I think you're confusing various different things as "neurosymbolic AI". There is a NeSy symposium and I happen to have met many of the people there, and they are not GOFAI ideologues, rather they recognise the obvious limitations of neural nets (i.e. they're crap at deduction, though great at induction) and…

I don't think symbolic approaches are completely useless. It's just that they're solving yesterday's problems 1.12% better. While ML is cracking open entirely new fields - and might go all the way to AGI, the way it's going now. One is near the end of its potential while another is only picking up steam. In many ways, the space ML dominates now is the space of "all the things symbolic approaches suck ass at". Which i…

Is SAT/SMT and theorem provers solving yesterday's problems 1.12% better?

Lots of the successes by LLMs that have been much celebrated rely on these.

Re: Executing programs inside transformers with exponentially faster inference

#132

Earlier quoted context omitted.

Full disclosure: all my published work is on symbolic machine learning (a.k.a. Inductive Logic Programming) :O I think you're confusing various different things as "neurosymbolic AI". There is a NeSy symposium and I happen to have met many of the people there, and they are not GOFAI ideologues, rather they recognise the obvious limitations of neural nets (i.e. they're crap at deduction, though great at induction) and…

I don't think symbolic approaches are completely useless. It's just that they're solving yesterday's problems 1.12% better. While ML is cracking open entirely new fields - and might go all the way to AGI, the way it's going now. One is near the end of its potential while another is only picking up steam. In many ways, the space ML dominates now is the space of "all the things symbolic approaches suck ass at". Which i…

Well, neural nets do what neural nets do best (not ML in general, which is a broader field), so if a lot of funding is going to neural nets then we'll see a lot of progress on the stuff neural nets are best suited for. No surprise. If Google et al were spending billions on symbolic AI maybe we'd see equally spectacular results there too. Maybe not. But we won't know because they don't.

There's no sense in which symbolic AI is at the end of its life and if you pay close attention you'll see that LLMs are trying to do all the things that symbolic AI is good at: major examples being reasoning, and planning from world models.

And as nextos says in the sibling comment most of the recent successes of LLMs in tasks that go beyond language generation, e.g. solving math olympiad problems, are the result of combining LLMs with symbolic verifiers.

>> While ML is cracking open entirely new fields - and might go all the way to AGI, the way it's going now.

I don't agree. Everything that neural nets do today, speech recognition, object identification in images, machine translation, language generation, program synthesis, game playing, protein folding, research automation, I mean every single thing really, is a task that comes from the depths of AI history. There's a big discussion to be had about why those tasks are "AI" tasks in the first place and what they have to do with "intelligence" in the broader sense (e.g. cats are intelligent but they can't generate any sort of text) but this discussion is constantly postponed as we all breathlessly run up the hill that neural nets are climbing. When we get to the top and find it was the wrong hill to climb, maybe we'll have that discussion at last, or maybe the entire industry, academia in tow, will run after the Next Big Thing in AI™ all over again. But- cracking open new fields? Nah. Not really.

AGI is not going to happen any time soon though. We have no idea what we're doing in terms of reproducing intelligence, that much is clear.

Re: Executing programs inside transformers with exponentially faster inference

#133
post #131

Earlier quoted context omitted.

I don't think symbolic approaches are completely useless. It's just that they're solving yesterday's problems 1.12% better. While ML is cracking open entirely new fields - and might go all the way to AGI, the way it's going now. One is near the end of its potential while another is only picking up steam. In many ways, the space ML dominates now is the space of "all the things symbolic approaches suck ass at". Which i…

Is SAT/SMT and theorem provers solving yesterday's problems 1.12% better? Lots of the successes by LLMs that have been much celebrated rely on these.

:waves:

Re: Executing programs inside transformers with exponentially faster inference

#134

Earlier quoted context omitted.

I don't think symbolic approaches are completely useless. It's just that they're solving yesterday's problems 1.12% better. While ML is cracking open entirely new fields - and might go all the way to AGI, the way it's going now. One is near the end of its potential while another is only picking up steam. In many ways, the space ML dominates now is the space of "all the things symbolic approaches suck ass at". Which i…

Well, neural nets do what neural nets do best (not ML in general, which is a broader field), so if a lot of funding is going to neural nets then we'll see a lot of progress on the stuff neural nets are best suited for. No surprise. If Google et al were spending billions on symbolic AI maybe we'd see equally spectacular results there too. Maybe not. But we won't know because they don't. There's no sense in which symbo…

The whole notion of "we need to know what intelligence is exactly to reproduce it" is completely and utterly wrong.

It's also the kind of thinking that results in "neurosymbolic garbage is good actually".

What neural nets do today is basically "everything humans do". There is no longer a list of "things computers can't do" - just a list of things computers do worse than the top 1% of humans. Ever shrinking.

Re: Executing programs inside transformers with exponentially faster inference

#135

Earlier quoted context omitted.

Well, for one, by eliminating external tool calling, the model gains an amount of security. This occurs because the tools being called by an LLM can be corrupted, and in this scenario corrupted tools would not be called.

Prompt injection is still a possibility, so while it improves the security posture, not by much.

Prompt injection will always be a possibility, it's a direct consequence of the fundamental nature being a fully general tool.

Re: Executing programs inside transformers with exponentially faster inference

#136

Earlier quoted context omitted.

Well, neural nets do what neural nets do best (not ML in general, which is a broader field), so if a lot of funding is going to neural nets then we'll see a lot of progress on the stuff neural nets are best suited for. No surprise. If Google et al were spending billions on symbolic AI maybe we'd see equally spectacular results there too. Maybe not. But we won't know because they don't. There's no sense in which symbo…

The whole notion of "we need to know what intelligence is exactly to reproduce it" is completely and utterly wrong. It's also the kind of thinking that results in "neurosymbolic garbage is good actually". What neural nets do today is basically "everything humans do". There is no longer a list of "things computers can't do" - just a list of things computers do worse than the top 1% of humans. Ever shrinking.

[deleted]

Re: Executing programs inside transformers with exponentially faster inference

#137

Earlier quoted context omitted.

Well, neural nets do what neural nets do best (not ML in general, which is a broader field), so if a lot of funding is going to neural nets then we'll see a lot of progress on the stuff neural nets are best suited for. No surprise. If Google et al were spending billions on symbolic AI maybe we'd see equally spectacular results there too. Maybe not. But we won't know because they don't. There's no sense in which symbo…

The whole notion of "we need to know what intelligence is exactly to reproduce it" is completely and utterly wrong. It's also the kind of thinking that results in "neurosymbolic garbage is good actually". What neural nets do today is basically "everything humans do". There is no longer a list of "things computers can't do" - just a list of things computers do worse than the top 1% of humans. Ever shrinking.

Well, for example a computer can't make me an omelette. There's tons of examples like that, pretty much everything humans "can do" with our bodies, that computers can't- not just because they don't have bodies, but because even when we give them bodies we can't program them to do the things we want them to. LLMs don't help at all here. They can easily fake knowing what to do but the -not few- attempts people have made to connect LLMs to a robot to get the LLM to drive the robot like a little AI brain have ... not really worked out? I guess? Not even self-driving cars use LLMs.

Speaking of self-driving cars' AIs, while they have plenty of machine learning components, e.g. for vision, SLAM, and so on, they are largely hand-coded, rule-based systems. Just like the good old days of GOFAI.

>> The whole notion of "we need to know what intelligence is exactly to reproduce it" is completely and utterly wrong.

Can you explain why it's completely wrong?

Re: Executing programs inside transformers with exponentially faster inference

#138

Earlier quoted context omitted.

Nope, they encoded or compiled in a simple VM / WASM interpreter to the transformer weights, there is no training. You'd be forgiven for this misreading, as they deliberately mislead early on that their model is (in principle) trainable, but later admit that their actual model is not actually differentiable, but that a differentiable approximation "should" still work (despite no info about what loss function or train…

Thanks, but where do they say that? I can only find this instance of "different" (as in "differentiable") in their article: Because the execution trace is part of the forward pass, the whole process remains differentiable: we can even propagate gradients through the computation itself. That makes this fundamentally different from an external tool. It becomes a trainable computational substrate that can be integrated…

In the section "Programs into weights & training beyond gradient descent", near the end, they say:

    [...] *the compilation machinery we built for generating those weights** can go further. In principle, arbitrary programs can be compiled directly into the transformer weights, bypassing the need to represent them as token sequences at all. [...] [my emphasis]
In the same section, they also continue:

    Weights become a deployment target: instead of learning software-like behavior, models contain compiled program logic.

    If logic can be compiled into weights, then gradient descent is no longer the only way to modify a model. Weight compilation provides another route for inserting structure, algorithms, and guarantees directly into a network.
So they (almost-invisibly) admit they compile in the weights, but make it clearer this was the whole intention the whole time in later sentences.

Re: Executing programs inside transformers with exponentially faster inference

#139

Earlier quoted context omitted.

The whole notion of "we need to know what intelligence is exactly to reproduce it" is completely and utterly wrong. It's also the kind of thinking that results in "neurosymbolic garbage is good actually". What neural nets do today is basically "everything humans do". There is no longer a list of "things computers can't do" - just a list of things computers do worse than the top 1% of humans. Ever shrinking.

Well, for example a computer can't make me an omelette. There's tons of examples like that, pretty much everything humans "can do" with our bodies, that computers can't- not just because they don't have bodies, but because even when we give them bodies we can't program them to do the things we want them to. LLMs don't help at all here. They can easily fake knowing what to do but the -not few- attempts people have mad…

> LLMs don't help at all here

You haven't heard about VLAs in robotics.

Post reply on HN