Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

471–478 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#471
post #370

Earlier quoted context omitted.

In my experience LLMs can't get basic western music theory right, there's no way I would use an LLM for something harder than that.

> In my experience LLMs can't get basic western music theory right, there's no way I would use an LLM for something harder than that. This take is completely oblivious, and frankly sounds like a desperate jab. There are a myriad of activities whose core requirement is a) derive info from a complex context which happens to be supported by a deep and plentiful corpus, b) employ glorified template and rule engines. LLMs…

A desperate jab? But I _want_ LLM's to be able to do basic, deterministic things accurately. Seems like I touched a nerve? Lol.

Re: Recent AI model progress feels mostly like bullshit

#472
I mean: * I don't think we've seen any major release or new architectural changes in the major (large companies) models recently

* Model creation has exploded with people training their own models and fine tunes, etc but these are all derivatives of parent models from large companies

So I'm not really sure what they mean when they refer to "recent model progress"...I don't think anybody is putting out a llama finetune saying "this is revolutionary!111" nor have I seen OAI, et al make any such claims either.

Is the sensation just because forward momentum is stalling while we wait for the next big leap?

Re: Recent AI model progress feels mostly like bullshit

#473

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

This is simply using LLMs directly. Google has demonstrated that this is not the way to go when it comes to solving math problems. AlphaProof, which used AlphaZero code, got a silver medal in last year's IMO. It also didn't use any human proofs(!), only theorem statements in lean, without their corresponding proofs [1].

[1] https://www.youtube.com/watch?v=zzXyPGEtseI

Re: Recent AI model progress feels mostly like bullshit

#474

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

Because of the vast number of problems reused, removing those data from training sets will just make models worse. Why would anyone do it?

Re: Recent AI model progress feels mostly like bullshit

#475

Earlier quoted context omitted.

It's obvious that humans imitate concepts and don't come up with things de-novo from a blank slate of pure intelligence. So your claim hinges on LLMs parrotting the words they are trained on. But they don't do that, their training makes them abstract over concepts and remix them in new ways to output sentences they weren't trained on, e.g.: Prompt: "Can you give me a URL with some novel components, please?" DuckDuckG…

> fireworks, cannons, jellyfish squeezing water out to accelerate, no sudies of orbits from moons and planets, no chemistry experiments, no inspiration from thousands of years of flamethrowers Fireworks, cannons, chemistry experiments and flamethrowers are all human inventions And yes, exactly! We studied orbits of moons and planets. We studied animals like Jellyfish. We choose to observe the world , we extracted dat…

Your point is contingent on sensor availability to an llm. Llms are a frozen human mind until they behave like live ml algos.

Re: Recent AI model progress feels mostly like bullshit

#476
post #154

Earlier quoted context omitted.

This is less an LLM thing than an information retrieval question. If you choose a model and tell it to “Search,” you find citation based analysis that discusses that he indeed had problems with alcohol. I do find it interesting it quibbles whether he was an alcoholic or not - it seems pretty clear from the rest that he was - but regardless. This is indicative of something crucial when placing LLMs into a toolkit. The…

I realise your answer wasn't assertive, but if I heard this from someone actively defending AI it would be a copout. If the selling point is that you can ask these AIs anything then one can't retroactively go "oh but not that" when a particular query doesn't pan out.

My point is the opposite of this point of view. I believe generative AI is the most significant advance since hypertext and the overlay of inferred semantic relationships via pagerank etc. In fact the creation of hypertext and the toolchains around it led to this point at all - neural networks were understood at that point and transformer attention is just an innovation. It’s the collective human assembly of language and visual interconnected knowledge at a pan cultural and global scale that enabled the current state.

The abilities of LLM alone to do astounding natural language processing beyond the ability of anything prior by unthinkable Turing test passing miles. The fact it can reason abductively, which computing techniques to date have been unable to is amazing. The fact you can mix it with multimodal regimes - images, motion, virtually anything that can be semantically linked via language, is breathtaking. The fact it can be augmented with prior computing techniques - IR, optimization, deductive solvers, and literally everything we’ve achieved to date should give anyone knowledgeable of such things shivers for what the future holds.

But I would never hold that generative AI techniques are replacements for known optimal techniques. But the ensemble is probably the solution to nearly every challenge we face. When we hit the limits of LLMs today, I think, well, at least we already have grand master beating chess solvers and it’s irrelevant the LLM can’t directly. The LLM and other generative AI techniques in my mind are like gasses that fill through learned approximation the things we’ve not been able to solve directly, including the assembly of those solutions ad hoc. This is why since the first time BERT came along I knew agent based techniques were the future.

Right now we live at time like early hypertext with respect to AI. Toolchains suck, LLMs are basically geocities pages with “under construction” signs. We will go through an explosive exploration, some stunning insights that’ll change the basic nature of our shared reality (some wonderful some insidious), then if we aren’t careful - and we rarely are - enshitification at scale unseen before.

Re: Recent AI model progress feels mostly like bullshit

#477

Earlier quoted context omitted.

This is less an LLM thing than an information retrieval question. If you choose a model and tell it to “Search,” you find citation based analysis that discusses that he indeed had problems with alcohol. I do find it interesting it quibbles whether he was an alcoholic or not - it seems pretty clear from the rest that he was - but regardless. This is indicative of something crucial when placing LLMs into a toolkit. The…

lotta words here to say AI can't do basic search right

Lotta words to say AI can’t do basic search in the same way a web browser can’t do basic search, but given a search engine both can.

Re: Recent AI model progress feels mostly like bullshit

#478

Earlier quoted context omitted.

lotta words here to say AI can't do basic search right

Lotta words to say AI can’t do basic search in the same way a web browser can’t do basic search, but given a search engine both can.

i don't know what this means
Post reply on HN