Live data from Hacker News

AI language models are struggling to “get” math

spectrum.ieee.org

111–120 of 201 posts

Re: AI language models are struggling to “get” math

#111

How much of this is just "AI is bad at everything", but in the math case, it's easier for the lay person to tell . It's all just passable garbled nonesense that the reader (goes to lengths) to interept based on their prior knowledge, which is not expressed in the syntax of what these systems output. In the case of mathematics, we're far less willing to "BS away" the interpretive failures. But if we were equally deman…

Language models aren't built for math. Their improvement/training cycles aren't sensitive to the exactness and rule-based nature of mathematical language, plus there are probably a lot of bad/misleading examples of math in the source data.

You'd have to be unrealistically pessimistic to call what GPT-3 and other huge language models produce "nonsense".

Re: AI language models are struggling to “get” math

#112
post #75

Earlier quoted context omitted.

> language models just need to translate problems into code of some kind that can be run to get the answer A huge "just"! Isn't this the magic step? Translating ambiguous symbols to meaning and combining them in meaningful ways is a big deal which, apparently, these AI models cannot do. They can just parrot things.

It’s already being done and will only get better: https://twitter.com/sergeykarayev/status/1569377881440276481

I suspect it's not solved, because solving this (beyond some trick/toy examples) is essentially solving General AI.

Re: AI language models are struggling to “get” math

#113

How much of this is just "AI is bad at everything", but in the math case, it's easier for the lay person to tell . It's all just passable garbled nonesense that the reader (goes to lengths) to interept based on their prior knowledge, which is not expressed in the syntax of what these systems output. In the case of mathematics, we're far less willing to "BS away" the interpretive failures. But if we were equally deman…

Language models aren't built for math. Their improvement/training cycles aren't sensitive to the exactness and rule-based nature of mathematical language, plus there are probably a lot of bad/misleading examples of math in the source data. You'd have to be unrealistically pessimistic to call what GPT-3 and other huge language models produce "nonsense".

It's not that they were not built for math, but more like verification is hard. But it's hard for humans as well. A large generative model + a fast verifier could do wonders.

AlphaGo was built on that - the model can propose moves, but you can verify who won in the end. There are some code generation models that write their own tests as well, or use externally provided tests to verify their solutions. The DeepMind matrix multiplication algorithm was also "learning from verification" of generated solutions, because it's trivial to do that. In general verification remains an open problem.

Re: AI language models are struggling to “get” math

#115

Earlier quoted context omitted.

This is factually wrong, both in terms of quantity and quality. Current AI models are not "just sort of repeating and copying from memory". This is just an incorrect characterization of how they work and how they perform. AI skeptics often say things like this then backpedal with something like "Well they aren't really repeating what they heard, but their generative model is just a slightly more sophisticated version…

can you actually share what "current AI models" are then? Not trying to be rude, but you just said "na ah" and then refused to argument any position.

Current LLMs are "modeling" something according to pretty much any sense of the word "model".

In the technical, computational linguistics sense, LLMs are language models that give a conditional posterior distribution over sentences. Given some (constrained) context, the model tells you the posterior distribution over sentences in or around that context.

In the nontechnical, layman sense of the word, they are a system that is used as an example of language. LLMs imitate language by generating new sentences. They are a "model" in the same way that an architectural model is a model, or in the same way that a statue is a model of a human.

The other point I disagreed with is the characterization that LLMs "just sort of repeat and copy from memory". I went into more detail about that in other replies.

Re: AI language models are struggling to “get” math

#116
post #97
post #80

Earlier quoted context omitted.

Indeed, there is lots of denial or ignorance in this thread (ignorance in the technical sense). AudioLM already produced impressive results and it's a tiny fraction of what is already possible because performance simply improves with scale. One can probably solve music generation today with a ~$1B budget for most purposes like film or game music, or personalized soundtracks. This is not science fiction.

I don't see a lot of progress in AudioLM compared to results from 2018: https://storage.googleapis.com/magentadata/papers/maestro/in... What's more interesting and concerning - listen carefully to the first piano continuation example from AudioLM, notice the similarity of the last 7 seconds to Moonlight sonata: https://youtu.be/4Tr0otuiQuU?t=516 I'm afraid we will see a lot of this with music generation models in the…

There are quite simple tricks to avoid repetition/copying in NNs, e.g. by (1) training a model to predict the "popularity" of the main model's outputs and penalizing popular/copied productions by backpropping through that model so as to decrease the predicted popularity, or (2) by conditioning on random inputs (LLMs can be prompted with imaginary "ID XXX" prefixes before each example to mitigate repetitions), or (3) by increasing temperature or optimizing for higher entropy. LLM outputs are already extremely diverse and verbatim copying is not a huge issue at all. The point being, all evidence points to this not being a show stopper if you massage these evolutionary methods for long enough in one or more of the various right ways.

Re: AI language models are struggling to “get” math

#117

How much of this is just "AI is bad at everything", but in the math case, it's easier for the lay person to tell . It's all just passable garbled nonesense that the reader (goes to lengths) to interept based on their prior knowledge, which is not expressed in the syntax of what these systems output. In the case of mathematics, we're far less willing to "BS away" the interpretive failures. But if we were equally deman…

> How much of this is just "AI is bad at everything", but in the math case, it's easier for the lay person to tell

Honestly, even as someone generally pretty dismissive of the AI hype, I'm not sure you can go that far. The whole reason we have specific mathematical notation is that human languages often are not super great at dealing with it, and English in particular is pretty abysmal for being both unambiguous and precise (and I'd be surprised if language models didn't end up suffering from biases analogous to how many image recognition AI models have been found to not deal well with a diverse set of human appearances). We don't teach math the same way we teach English, and we certainly don't expect people to be experts at teaching both, so why would we expect an AI model designed for language to be able to do math?

Re: AI language models are struggling to “get” math

#118
post #84

Earlier quoted context omitted.

Have you ever used Github CoPilot? It does a lot of useful work, automating away rote typing in programming. Have you tried Dall-E or Stable Diffusion? They make good looking images. This comment seems completely unmoored from where the state of the art is right now.

sure, but co-pilot is mostly just copying code (see, for example, the issue with it producing quake source code). If you think of AI as a dial from sample(data) to mean(data), then as the dial is turned towards the mean() you get more "generic" results, but also more garbled ones. Copilot is more like a search engine, having turned the dial more towards sample(). The real invention of the NN is simply to provide that…

Even if AI got to the point of perfectly passing every expert-level Turing test your degree of rigor as to what "thinking" is would never truly permit any belief of AI having struck the golden nugget of intelligence.

Imagine if we were all self-replicating computers, and certain members of this silicon race began experimenting with making creatures with carbon macro-molecules to create organic intelligence, you could make the same claim in the other direction:

"There has been no step-change advancement in Organic Intelligence in, perhaps, 50 years. All we see today is a product of cell count, in neurotransmitter chemistry able to compress TBs of experiences into c. 300B neurons."

Re: AI language models are struggling to “get” math

#119
It is a great sign that we are building AI in the right direction. Before building artificial human intelligence, it makes sense to get to the intelligence level of a mosquito or fly, then go to more intelligent animals in later iterations.

As most of the human knowledge is encoded in videos, getting better at understanding / generating videos will clearly get us closer to make computers understand the world.

Re: AI language models are struggling to “get” math

#120
post #84

How much of this is just "AI is bad at everything", but in the math case, it's easier for the lay person to tell . It's all just passable garbled nonesense that the reader (goes to lengths) to interept based on their prior knowledge, which is not expressed in the syntax of what these systems output. In the case of mathematics, we're far less willing to "BS away" the interpretive failures. But if we were equally deman…

Have you ever used Github CoPilot? It does a lot of useful work, automating away rote typing in programming. Have you tried Dall-E or Stable Diffusion? They make good looking images. This comment seems completely unmoored from where the state of the art is right now.

I agree. It's possible to point out the clear limitations of current AI without being oblivious to the huge, indisputable advances that have occurred.

People thought it might take centuries for a computer to defeat a top human in Go. Then deep learning showed up and a few years later it's the opposite.

A lot of the things deep learning methods are doing now are things no one had any idea how long research would take to achieve, or if they were even possible.

Personally, I think we are currently hitting some walls that might take a while to climb before we get to AGI, but I am very impressed at the recent progress.

Post reply on HN