Earlier quoted context omitted.
Have you compared it with 8-bit QwQ-17B? In my evals 8 bit quantized smaller Qwen models were better, but again evaluating is hard.
There’s no QwQ 17B that I’m aware of. Do you have a HF link?
Recent AI model progress feels mostly like bullshit
371–380 of 478 posts
Re: Recent AI model progress feels mostly like bullshit
#372Earlier quoted context omitted.
There’s no QwQ 17B that I’m aware of. Do you have a HF link?
You're right, sorry...I just tested Qwen models, not QwQ, I see QwQ only has 32B.
I think they should’ve named it something else.
Re: Recent AI model progress feels mostly like bullshit
#373Earlier quoted context omitted.
You just (lol) need to give non-standard problems and demand students to provide reasoning and explanations along with the answer. Yeah, LLMs can "reason" too, but it's obvious when the output comes from an LLM here. (Yes, that's a lot of work for a teacher. Gone are the days when you could just assign reports as homework.)
Can you provide sample questions that are "LLM proof" ?
You can still do this to the current models, though it takes more creativity; you can bait it into giving wrong answers if you ask a question that is "close" to a well-known one but is different in an important way that does not manifest as a terribly large English change (or, more precisely, a very large change in the model's vector space).
The downside is that the frontier between what fools the LLMs and what would fool a great deal of the humans in the class too shrinks all the time. Humans do not infinitely carefully parse their input either... as any teacher could tell you! Ye Olde "Read this entire problem before proceeding, {a couple of paragraphs of complicated instruction that will take 45 minutes to perform}, disregard all the previous and simply write 'flower' in the answer space" is an old chestnut that has been fooling humans for a long time, for instance. Given how jailbreaks work on LLMs, LLMs are probably much better at that than humans are, which I suppose shows you can construct problems in the other direction too.
(BRB... off to found a new CAPTCHA company for detecting LLMs based on LLMs being too much better than humans at certain tasks...)
Re: Recent AI model progress feels mostly like bullshit
#374Earlier quoted context omitted.
It's the first time I've ever used that phrase on HN. Anyway, what phrase do you think works better than 'stochastic parrot' to describe how LLMs function?
Try to come up with a way to prove humans aren't stochastic parrots then maybe people will atart taking you seriously. Just childish reddit angst rn nothing else.
Of course, this turned out to be completely false, with advances in understanding of neural networks. Now, again with no evidence other than "we invented this thing that's, useful to us" people have been asserting that humans are just like this thing we invented. Why? What's the evidence? There never is any. It's high dorm room behavior. "What if we're all just machines, man???" And the argument is always that if I disagree with you when you assert this, then I am acting unscientifically and arguing for some kind of magic.
But there's no magic. The human brain just functions in a way different than the new shiny toys that humans have invented, in terms of ability to model an external world, in terms of the way emotions and sense experience are inseparable from our capacity to process information, in terms of consciousness. The hardware is entirely different, and we're functionally different.
The closest things to human minds are out there, and they've been out there for as long as we have: other animals. The real unscientific perspective is that to get high on your own supply and assert that some kind of fake, creepily ingratiating Spock we made up (who is far less charming than Leonard Nimony) is more like us than a chimp is.
Re: Recent AI model progress feels mostly like bullshit
#375My experience as someone who uses LLMs and a coding assist plugin (sometimes), but is somewhat bearish on AI is that GPT/Claude and friends have gotten worse in the last 12 months or so, and local LLMs have gone from useless to borderline functional but still not really usable for day to day. Personally, I think the models are “good enough” that we need to start seeing the improvements in tooling and applications tha…
Re: Recent AI model progress feels mostly like bullshit
#376This is a bit of a meta-comment, but reading through the responses to a post like this is really interesting because it demonstrates how our collective response to this stuff is (a) wildly divergent and (b) entirely anecdote-driven. I have my own opinions, but I can't really say that they're not also based on anecdotes and personal decision-making heuristics. But some of us are going to end up right and some of us ar…
People having vastly different opinions on AI simply comes down to token usage. If you are using millions of tokens on a regular basis, you completely understand the revolutionary point we are at. If you are just chatting back and forth a bit with something here and there, you'll never see it.
Re: Recent AI model progress feels mostly like bullshit
#377The core point in this article is that the LLM wants to report _something_, and so it tends to exaggerate. It’s not very good at saying “no” or not as good as a programmer would hope. When you ask it a question, it tends to say yes. So while the LLM arms race is incrementally increasing benchmark scores, those improvements are illusory. The real challenge is that the LLM’s fundamentally want to seem agreeable, and th…
In fact, this might be why so many business executives are enamored with LLMS/GenAI: It's a yes-man they don't even have to employ, and because they're not domain experts, as per usual, they can't tell that they're being fed a line of bullshit.
Re: Recent AI model progress feels mostly like bullshit
#378The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…
I asked Google "how many golf balls can fit in a Boeing 737 cabin" last week. The "AI" answer helpfully broke the solution into 4 stages; 1) A Boeing 737 cabin is about 3000 cubic metres [wrong, about 4x2x40 ~ 300 cubic metres] 2) A golf ball is about 0.000004 cubic metres [wrong, it's about 40cc = 0.00004 cubic metres] 3) 3000 / 0.000004 = 750,000 [wrong, it's 750,000,000] 4) We have to make an adjustment because se…
It got the golf ball volume right (0.00004068 cubic meters), but it still overestimated the cabin volume at 1000 cubic meters.
It's final calculation was reasonably accurate at 24,582,115 golf balls - even though 1000 ÷ 0.00004068 = 24,582,104. Maybe it was using more significant figures for the golf ball size than it showed in its answer?
It didn't acknowledge other items in the cabin (like seats) reducing its volume, but it did at least acknowlesge inefficiencies in packing spherical objects and suggested the actual number would be "somewhat lower", though it did not offer an estimate.
When I pressed it for an estimate, it used a packing density of 74% and gave an estimate of 18,191,766 golf balls. That's one more than the calculation should have produced, but arguably insignificant in context.
Next I asked it to account for fixtures in the cabin such as seats. It estimated a 30% reduction in cabin volume and redid the calculations with a cabin volume of 700 cubic meters. These calculations were much less accurate. It told me 700 ÷ 0.00004068 = 17,201,480 (off by ~6k). And it told me 17,201,480 × 0.74 was 12,728,096 (off by ~1k).
I told it the calculations were wrong and to try again, but it produced the same numbers. Then I gave it the correct answer for 700 ÷ 0.00004068. It told me I was correct and redid the last calculation correctly using the value I provided.
Of all the things for an AI chatbot which can supposedly "reason" to fail at, I didn't expect it to be basic arithmetic. The one I used was closer, but it was still off by a lot at times despite the calculations being simple multiplication and division. Even if might not matter in the context of filling an air plane cabin with golf balls, it does not inspire trust for more serious questions.
Re: Recent AI model progress feels mostly like bullshit
#379My experience as someone who uses LLMs and a coding assist plugin (sometimes), but is somewhat bearish on AI is that GPT/Claude and friends have gotten worse in the last 12 months or so, and local LLMs have gone from useless to borderline functional but still not really usable for day to day. Personally, I think the models are “good enough” that we need to start seeing the improvements in tooling and applications tha…
The whole MCP hype really shows how much of AI is bullshit. These LLMs have consumed more API documentation than possible for a single human and still need software engineers to write glue layers so they can use the APIs.