Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

1–10 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#3
I'd say most of the recent AI model progress has been on price.

A 4-bit quant of QwQ-32B is surprisingly close to Claude 3.5 in coding performance. But it's small enough to run on a consumer GPU, which means deployment price is now down to $0.10 per hour. (from $12+ for models requiring 8x H100)

Re: Recent AI model progress feels mostly like bullshit

#6
> So maybe there's no mystery: The AI lab companies are lying, and when they improve benchmark results it's because they have seen the answers before and are writing them down. [...then says maybe not...]

Well.. they've been caught again and again red handed doing exactly this. Fool me once shame on you, fool me 100 times shame on me.

Re: Recent AI model progress feels mostly like bullshit

#7
> Since 3.5-sonnet, we have been monitoring AI model announcements, and trying pretty much every major new release that claims some sort of improvement. Unexpectedly by me, aside from a minor bump with 3.6 and an even smaller bump with 3.7, literally none of the new models we've tried have made a significant difference on either our internal benchmarks or in our developers' ability to find new bugs. This includes the new test-time OpenAI models.

This is likely a manifestation of the bitter lesson[1], specifically this part:

> The ultimate reason for this is Moore's law, or rather its generalization of continued exponentially falling cost per unit of computation. Most AI research has been conducted as if the computation available to the agent were constant (in which case leveraging human knowledge would be one of the only ways to improve performance) but, over a slightly longer time than a typical research project [like an incremental model update], massively more computation inevitably becomes available.

(Emphasis mine.)

Since the ultimate success strategy of the scruffies[2] or proponents of search and learning strategies in AI is Moore's Law, short term gains using these strategies will be miniscule. It is over at least a five year period that their gains will be felt the most. The neats win the day in the short term, but the hare in this race will ultimately give away to the steady plod of the tortoise.

1: http://www.incompleteideas.net/IncIdeas/BitterLesson.html

2: https://en.m.wikipedia.org/wiki/Neats_and_scruffies#CITEREFM...

Re: Recent AI model progress feels mostly like bullshit

#10
post #5

This was published the day before Gemini 2.5 was released. I'd be interested if they see any difference with that model. Anecdotally, that is the first model that really made me go wow and made a big difference for my productivity.

I doubt it. It still flails miserably like the other models on anything remotely hard, even with plenty of human coaxing. For example, try to get it to solve: https://www.janestreet.com/puzzles/hall-of-mirrors-3-index/
Post reply on HN