Earlier quoted context omitted.
Massive search overlap though - and some questions (like the golf ball puzzle) can be cached for a long time.
AFAIK they got 15% of unseen queries everyday, so it might be not very simple to design an effective cache layer on that. Semantic-aware clustering of natural language queries and projecting them into a cache-able low rank dimension is a non-trivial problem. Of course, LLM can effectively solve that, but then what's the point of using cache when you need LLM for clustering queries...
Recent AI model progress feels mostly like bullshit
441–450 of 478 posts
Re: Recent AI model progress feels mostly like bullshit
#442Earlier quoted context omitted.
Yeah I’m a computational biology researcher. I’m working on a novel machine learning approach to inferring cellular behavior. I’m currently stumped why my algorithm won’t converge. So, I describe the mathematics to ChatGPT-o3-mini-high to try to help reason about what’s going on. It was almost completely useless. Like blog-slop “intro to ML” solutions and ideas. It ignores all the mathematical context, and zeros in o…
I tend to prefer Claude over all things ChatGPT so maybe give the latest model a try -- although in some way I feel like 3.7 is a step down from the prior 3.5 model
Re: Recent AI model progress feels mostly like bullshit
#443Earlier quoted context omitted.
My point wasn't chess specific or that they couldn't have specific training for it. It was a more general "here is something that LLMs clearly aren't being trained for currently, but would also be solvable through reasoning skills" Much in the same way a human who only just learnt the rules but 0 strategy would very, very rarely lose here These companies are shouting that their products are passing incredibly hard ex…
>If an LLM is trained for chess then its performance would just come from memorization, not any kind of "reasoning". If you think you can play chess at that level over that many games and moves with memorization then i don't know what to tell you except that you're wrong. It's not possible so let's just get that out of the way. >These companies are shouting that their products are passing incredibly hard exams, solvi…
Yes, that's all there is to it and it's not enough. I ain't paying for another defective organism that makes mistakes in entirely novel ways. At least with humans you know how to guide them back on course.
If that's the peak of "AI" evolution today, I am not impressed.
Re: Recent AI model progress feels mostly like bullshit
#444Earlier quoted context omitted.
>The whole premise on which the immense valuations of these AI companies is based on is that they are learning general reasoning skills from their training on language. And they do, just not always in the ways we expect. >This whole premise crashes and burns if you need task-specific training, like explicit chess training. Everyone needs task specific training. Any human good at chess or anything enough to make it a…
Chess is a very simple game, and having basic general reasoning skills is more than enough to learn how to play it. It's not some advanced mathematics or complicated human interaction - it's a game with 30 or so fixed rules. And chess manuals have numerous examples of actual chess games, it's not like they are pure text talking about the game. So, the fact that LLMs can't learn this sample game despite probably inclu…
Re: Recent AI model progress feels mostly like bullshit
#445There are real and obvious improvements in the past few model updates and I'm not sure what the disconnect there is. Maybe it's that I do have PhD level questions to ask them, and they've gotten much better at it. But I suspect that these anecdotes are driven by something else. Perhaps people found a workable prompt strategy by trial and error on an earlier model and it works less well with later models. Or perhaps t…
Why? Because I know so little about chemistry myself that I wouldn't even know what to start asking the model as to be impressed by the answer.
For the model to be useful at all, I would have to learn basic chemistry myself.
Many though I suspect are in this same situation with all subjects. They really don't know much of anything and are therefore unimpressed by the models response in the same way I am not impressed with chemistry responses.
Re: Recent AI model progress feels mostly like bullshit
#446This is a bit of a meta-comment, but reading through the responses to a post like this is really interesting because it demonstrates how our collective response to this stuff is (a) wildly divergent and (b) entirely anecdote-driven. I have my own opinions, but I can't really say that they're not also based on anecdotes and personal decision-making heuristics. But some of us are going to end up right and some of us ar…
You want to block subjectivity? Write some formulas. There are three questions to consider: a) Have we, without any reasonable doubt, hit a wall for AI development? Emphasis on "reasonable doubt". There is no reasonable doubt that the Earth is roughly spherical. That level of certainty. b) Depending on your answer for (a), the next question to consider is if we the humans have motivations to continue developing AI. c…
Re: Recent AI model progress feels mostly like bullshit
#447Earlier quoted context omitted.
Could you share any successful prompting techniques for grounding 3.7, even just a project-specific example?
I use this: I don't want to drastically change my current code, nor do I like being told to create several new files and numerous functions/classes to solve this problem. I want you to think clearly and be focused on the task and don't get wild! I want the most straightforward approach which is elegant, intuitive, and rock solid.
Not cool, Claude 3.7, not cool.
Re: Recent AI model progress feels mostly like bullshit
#448Earlier quoted context omitted.
What are you, an LLM? Look at the results of the first twenty hits and come back, then tell me that they don't speak to that specific issue.
Widely reported does not imply widely known.
Re: Recent AI model progress feels mostly like bullshit
#449Earlier quoted context omitted.
How do you make homework assignments LLM-proof? There may be a huge business opportunity if that actually works, because LLMs are destroying education at a rapid pace.
By giving pen and paper exams and telling your students that the only viable preparation strategy is doing the hw assignments themselves :)