Earlier quoted context omitted.
Yes, here's the link: https://arxiv.org/abs/2503.21934v1 Anecdotally, I've been playing around with o3-mini on undergraduate math questions: it is much better at "plug-and-chug" proofs than GPT-4, but those problems aren't independently interesting, they are explicitly pedagogical. For anything requiring insight, it's either: 1) A very good answer that reveals the LLM has seen the problem before (e.g. naming the theo…
This is a paper by INSAIT researchers - a very young institute which hired most of its PHD staff only in the last 2 years, basically onboarding anyone who wanted to be part of it. They were waiving their BG-GPT on national TV in the country as a major breakthrough, while it was basically was a Mistral fine-tuned model, that was eventually never released to the public, nor the training set. Not sure whether their (INS…
Recent AI model progress feels mostly like bullshit
411–420 of 478 posts
Re: Recent AI model progress feels mostly like bullshit
#412Earlier quoted context omitted.
3.7 is like a wild horse. you really must ground it with clear instructions. it sucks that it doesn't automatically know that but it's tameable.
Could you share any successful prompting techniques for grounding 3.7, even just a project-specific example?
I don't want to drastically change my current code, nor do I like being told to create several new files and numerous functions/classes to solve this problem. I want you to think clearly and be focused on the task and don't get wild! I want the most straightforward approach which is elegant, intuitive, and rock solid.Re: Recent AI model progress feels mostly like bullshit
#413Earlier quoted context omitted.
I would go even further than TFA. In my personal experience using Windsurf daily, Sonnet 3.5 is still my preferred model. 3.7 makes many more changes that I did not ask for, often breaking things. This is an issue with many models, but it got worse with 3.7.
Yea, I've experienced this too with 3.7. Not always though. It has been helpful for me more often than not helpful. But yea 3.5 "felt" better to me. Part of me thinks this is because I expected less of 3.5 and therefore interacted with it differently. It's funny because it's unlikely that everyone interacts with these models in the same way. And that's pretty much guaranteed to give different results. Would be intere…
This would be so useful. I have thought about this missing piece a lot.
Different tools like Cursor vs. Windsurf likely have their own system prompts for each model, so the testing really needs to be done in the context of each tool.
This seems somewhat straightforward to do using a testing tool like Playwright, correct? Whoever first does this successfully with have a popular blog/site on their hands.
Re: Recent AI model progress feels mostly like bullshit
#414Earlier quoted context omitted.
And then within a week, Gemini 2.5 was tested and got 25%. Point is AI is getting stronger. And this only suggested LLMs aren't trained well to write formal math proofs, which is true.
> within a week How do we know that Gemini 2.5 wasn't specifically trained or fine-tuned with the new questions? I don't buy that a new model could suddenly score 5 times better than the previous state-of-the-art models.
And to be clear, that's pretty much all this was: there's six problems, it got almost-full credit on one and half credit on another and bombed the rest, whereas all the other models bombed all the problems.
Re: Recent AI model progress feels mostly like bullshit
#415Earlier quoted context omitted.
Bold claim! Let's see what that 25% is. I guarantee it is the portion of the exam which is trivially answerable if you have a stored database of all previous math exams ever written to consult.
There is 0% of the exam which is trivially answerable. The entire point of USAMO problems is that they demand novel insight and rigorous, original proofs. They are intentionally designed not to be variations of things you can just look up. You have to reason your way through, step by logical step. Getting 25% (~11 points) is exceptionally difficult. That often means fully solving one problem and maybe getting solid p…
That's true, but of course, not what I claimed.
The claim is that, given the ability to memorize an every mathematical result that has ever been published (in print or online), it is not so difficult to get 25% correct on an exam by pattern matching.
Note that this is skill is, by definition, completely out of the reach of any human being, but that possessing it does not imply creativity or the ability to "think".
Re: Recent AI model progress feels mostly like bullshit
#416Earlier quoted context omitted.
I think there is a big divide here. Every adult on earth knows magic is "fake", but some can still be amazed and entertained by it, while others find it utterly boring because it's fake, and the only possible (mildly) interesting thing about it is to try to figure out what the trick is. I'm in the second camp but find it kind of sad and often envy the people who can stay entertained even though they know better.
I think magic is extremely interesting (particularly close-up magic), but I also hate the mindset (which seems to be common though not ubiquitous) that stigmatizes any curiosity in how the trick works. In my view, the trick as it is intended to appear to the audience and the explanation of how the trick is performed are equal and inseparable aspects of my interest as a viewer. Either one without the other is less int…
As a long-time close-up magician and magical inventor who's spent a lot of time studying magic theory (which has been a serious field of magical research since the 1960s), it depends on which way we interpret "how the trick works." Frankly, for most magic tricks the method isn't very interesting, although there are some notable exceptions where the method is fascinating, sometimes to the extent it can be far more interesting than the effect it creates.
However, in general, most magic theorists and inventors agree that the method, for example, "palm a second coin in the other hand", isn't usually especially interesting. Often the actual immediate 'secret' of the method is so simple and, in hindsight, obvious that many non-magicians feel rather let down if the method is revealed. This is the main reason magicians usually don't reveal secret methods to non-magicians. It's not because of some code of honor, it's simply because the vast majority of people think they'll be happy if they know the secret but are instead disappointed.
Where studying close-up magic gets really fascinating is understanding why that simple, obvious thing works to mislead and then surprise audiences in the context of this trick. Very often changing subtle things seemingly unrelated to the direct method will cause the trick to stop fooling people or to be much less effective. Comparing a master magician to even a competent, well-practiced novice performing the exact same effect with the same method can be a night and day difference. Typically, both performances will fool and entertain audiences but the master's performance can have an intensely more powerful impact. Like leaving most audience members in stunned shock vs just pleasantly surprised and fooled. While neither the master nor novice's audiences have any idea of the secret method, this dramatic difference in impact is fascinating because careful deconstruction reveals it often has little to do with mechanical proficiency in executing the direct method. In other words, it's rarely driven by being able to do the sleight of hand faster or more dexterously. I've seen legendary close-up masters like a Dai Vernon or Albert Goshman when in their 80s and 90s perform sleight of hand with shriveled, arthritic hands incapable of even cleanly executing a basic palm, absolutely blow away a roomful of experienced magicians with a trick all the magicians already knew. How? It turns out there's something deep and incredibly interesting about the subtle timing, pacing, body language, posture, and psychology surrounding the "secret method" that elevates the impact to almost transcendence compared to a good, competent but uninspired performance of the same method and effect.
Highly skilled, experienced magicians refer to the complex set of these non-method aspects, which can so powerfully elevate an effect to another level, as "the real work" of the trick. At the top levels, most magicians don't really care about the direct methods which some audience members get so obsessed about. They aren't even interesting. And, contrary to what most non-magicians think, these non-methods are the "secrets" master magicians tend to guard from widespread exposure. And it's pretty easy to keep this crucially important "real work" secret because it's so seemingly boring and entirely unlike what people expect a magic secret to be. You have to really "get it" on a deeper level to even understand that what elevated the effect was intentionally establishing a completely natural-seeming, apparently random three-beat pattern of motion and then carefully injecting a subtle pause and slight shift in posture to the left six seconds before doing "the move". Audiences mistakenly think that "the hidden move" is the secret to the trick when it's just the proximate first-order secret. Knowing that secret won't get you very far toward recreating the absolute gob-smacking impact resulting from a master's years of experimentation figuring out and deeply understanding which elements beyond the "secret method" really elevate the visceral impact of the effect to another level.
Re: Recent AI model progress feels mostly like bullshit
#417Earlier quoted context omitted.
> I have seen much AI output which is extraordinary it's funny how one serious fail can impact my point of view so dramatically. I feel the same way. It's like discovering for the first time that magicians aren't doing "real" magic, just sleight of hand and psychological tricks. From that point on, it's impossible to be convinced that a future trick is real magic, no matter how impressive it seems. You know it's fake…
I think there is a big divide here. Every adult on earth knows magic is "fake", but some can still be amazed and entertained by it, while others find it utterly boring because it's fake, and the only possible (mildly) interesting thing about it is to try to figure out what the trick is. I'm in the second camp but find it kind of sad and often envy the people who can stay entertained even though they know better.
I look at optical illusions like The Dress™ and am impressed that I cannot force my brain to see it correctly even though I logically know what color it is supposed to be.
Finding new ways that our brains can be fooled despite knowing better is kind of a fun exercise in itself.
Re: Recent AI model progress feels mostly like bullshit
#418Earlier quoted context omitted.
How do you make homework assignments LLM-proof? There may be a huge business opportunity if that actually works, because LLMs are destroying education at a rapid pace.
By giving pen and paper exams and telling your students that the only viable preparation strategy is doing the hw assignments themselves :)
After all, they will grow up next to these things. They will do the homework today, by the time they graduate the LLM will take their job. There might be human large langage model managers for a while, soon to be replaced by the age of idea men.
Re: Recent AI model progress feels mostly like bullshit
#419Earlier quoted context omitted.
I never said it was sustainable, and even if it was, OP asked for a business model. Customers don’t need a business model, they’re customers. The same is true for any non essential good or service.
Than any silly idea can be a business model. Suppose I collect dust from my attic and hope to sell it as an add-on on my neighbor's lemonade stand, with a hefty profit for the neighbor, who is getting paid by me $10 to add a handful of dust in each glass and sell it to the customers for $1. The neighbor accepts. It's a business model, at least until I don't run of existing funds or the last customer leaves in disguis…
Indeed it can. The difference between a business model and a viable business model is one word - viable.
If you asked me 18 years ago was "giving away a video game and selling cosmetics" a viable business model I would have laughed at you.If you asked me in 2019 I would probably give you money. If you asked me in 2025, I'd probably laugh at you again.
> and I need to borrow larger an larger lumps of money each time to keep spinning the wheel...
Or you figure out a way to to sell it to your neighbour for $0.50 and he can sell it on for $1.
The play is clear at every level - Nvidia Sell GPUs, OpenAI sell models, and SAAS sell prompts + UI's. Whether or not any of them are viable remains to be seen. Personally, I wouldn't take the bet.
Re: Recent AI model progress feels mostly like bullshit
#420Earlier quoted context omitted.
I don't think I agree, entirely. The problem is that up until _very_ recently, it's been possible to get LLMs to generate interesting and exciting results (as a result of all the API documentation and codebases they've inhaled), but it's been very hard to make that usable. I think we need to be able to control the output format of the LLMs in a better way before we can work on what's in the output. I don't konw if MC…
That's reasonable along with your comment below too, but when you have the ceo of anthropic saying "AI will write all code for software engineers within a year" last month I would say that is pretty hard to believe given how it performs without user intervention (MCP etc...). It feels like bullshit just like the self driving car stuff did ~10 years ago.