Earlier quoted context omitted.
It's worth pointing out that GPT-4.5 seems focused on better pre-training and doesn't include reasoning. I think GPT-5 - if/when it happens - will be 4.5 with reasoning, and as such it will feel very different. The barrier, is the computational cost of it. Once 4.5 gets down to similar costs to 4.0 - which could be achieved through various optimization steps (what happened to the ternary stuff that was published last…
Is it fair to still call LLMs stochastic parrots now that they are enriched with reasoning? Seems to me that the simple procedure of large-scale sampling + filtering makes it immediately plausible to get something better than the training distribution out of the LLM. In that sense the parrot metaphor seems suddenly wrong. I don’t feel like this binary shift is adequately accounted for among the LLM cynics.
GPT-4.5
891–900 of 1001 posts
Re: GPT-4.5
#892Earlier quoted context omitted.
I'm not convinced that LLMs in their current state are really making anyone's lives much better though. We really need more research applications for this technology for that to become apparent. Polluting the internet with regurgitated garbage produced by a chat bot does not benefit the world. Increasing the productivity of software developers does not help to the world. Solving more important problems should be the…
LLM's have been extremely useful for me. They are incredibly powerful programmers, from the perspective of people who aren't programmers . Just this past week claude 3.7 wrote a program for us to use to quickly modernize ancient (1990's) proprietary manufacturing machine files to contemporary automation files. This allowed us to forgo a $1k/yr/user proprietary software package that would be able to do the same. The p…
Eg as i've been trying Claude Code i still feel the need to babysit it with my primary work, and so i'd rather do it myself. However while i'm working if it could sit there and monitor it, note fixes, tests and documentation and then stub them in during breaks i think there's a lot of time savings to be gained.
Ie keep the doing simple tasks that it can get right 99% of the time and get it out of the way.
I also suspect there's context to be gained in watching the human work. Not learning per say, but understanding the areas being worked on, improving intuition on things the human needs or cares about, etc.
A `cargo lint --fix` on steroids is "simple" but still really sexy imo.
Re: GPT-4.5
#893Earlier quoted context omitted.
Didn't seem to realize that "Still more coherent than the OpenAI lineup" wouldn't make sense out of context. (The actual comment quoted there is responding to someone who says they'd name their models Foo, Bar, Baz.)
Wonder if there’s some pro-OpenAI system prompt getting in the way of that.
Re: GPT-4.5
#894Earlier quoted context omitted.
I agree unfortunately. I might be a bit of an extremist on this issue. I genuinely think that building agentic ASI is suicidally stupid and we just shouldn’t do it. All the utopian visions we hear from the optimists describe unstable outcomes. A world populated by super-intelligent agents will be incredibly dangerous even if it appears initially to have gone well. We’ll have built a paradise in which we can never rel…
What's the difference between your "agentic AIs" and, say, "script kiddies" or "expert anarchist/black-hat hackers"? It's been obvious for a while that the narrow-waist APIs between things matter, and apparent that agentic AI is leaning into adaptive API consumption, but I don't see how that gives the agentic client some super-power we don't already need to defend against since before AGI we already have HGI (human g…
Intelligence. I'm talking about super-intelligence. If you want to know what it feels like to be intellectually outclassed by a machine, download the latest Go engine and have fun losing again and again while not understanding why. Now imagine an ASI that isn't confined to the Go board, but operating out in the world. It's doing things you don't like at speeds you can scarcely comprehend and there's not a thing you can do about it.
Re: GPT-4.5
#895Earlier quoted context omitted.
Is it fair to still call LLMs stochastic parrots now that they are enriched with reasoning? Seems to me that the simple procedure of large-scale sampling + filtering makes it immediately plausible to get something better than the training distribution out of the LLM. In that sense the parrot metaphor seems suddenly wrong. I don’t feel like this binary shift is adequately accounted for among the LLM cynics.
it was never fair to call them stochastic parrots and anybody who is paying any attention knows that sequence models can generalize at least partially OOD
Re: GPT-4.5
#896Earlier quoted context omitted.
I'm not convinced that LLMs in their current state are really making anyone's lives much better though. We really need more research applications for this technology for that to become apparent. Polluting the internet with regurgitated garbage produced by a chat bot does not benefit the world. Increasing the productivity of software developers does not help to the world. Solving more important problems should be the…
LLM's have been extremely useful for me. They are incredibly powerful programmers, from the perspective of people who aren't programmers . Just this past week claude 3.7 wrote a program for us to use to quickly modernize ancient (1990's) proprietary manufacturing machine files to contemporary automation files. This allowed us to forgo a $1k/yr/user proprietary software package that would be able to do the same. The p…
Re: GPT-4.5
#897Just going to put this here: https://www.wheresyoured.at/wheres-the-money/
Good write-up.
But it focuses too much on the big companies. Many indiehackers have figured out how to make profit with AI:
1. No free tier. Just provide a good landing page.
2. Ship fast. Ship iteratively. Employ no one besides yourself.
3. Profit.
The old silicon valley idea that you need to raise a bunch of money, hire a bunch of devs, and scale a ton to satisfy investors is dying rapidly for software. You can code and profit millions as just a single person company, especially in the age of cursor.
Re: GPT-4.5
#898The results for GPT - 4.5 are in for Kagi LLM benchmark too. It does crush our benchmark - time to make new? ;) - with performance similar of that of reasoning models. It does come at a great price both in cost and speed. A monster is what they created. But looking at the tasks it fails, some of them my 9 year old would solve. Still in this weird limbo space of super knowledge and low intelligence. May be remembered…
If Gemini 2 is the top in your benchmark, make sure to re-check your benchmark.
Re: GPT-4.5
#899A bit better at coding than ChatGPT 4o but not better than o3-mini - there is a chart near the bottom of the page that is easy to overlook: - ChatGPT 4.5 on AWS Bench verified: 38.0% - ChatGPT 4o on AWS Bench verified: 30.7% - OpenAI o3-mini on AWS Bench verified: 61.0% BTW Anthropic Claude 3.7 is better than o3-mini at coding at around 62-70% [1]. This means that I'll stick with Claude 3.7 for the time being for my…
I don't see Claude 3.7 on the official leaderboard. The top performer on the leaderboard right now is o1 with a scaffold (W&B Programmer O1 crosscheck5) at 64.6%: https://www.swebench.com/#verified . If Claude 3.7 achieves 70.3%, it's quite impressive, it's not far from 71.7% claimed by o3, at (presumably) much, much lower costs.
Re: GPT-4.5
#900Earlier quoted context omitted.
I don't know if I fully agree. The input clearly shows the need for emotional support more than "how do I pass this test?" The answer by 4o is comical even if you know you're talking to a machine. It reminds me of the advice to "not offer solutions when a woman talks about her problems, but just listen."
How could a machine provide emotional support? When I ask questions like this to LLMs, it's always to brainstorm solutions. I get annoyed when I receive fake-attention follow-up questions instead. I guess there's a trade-off between being human and being useful. But this isn't unique to LLMs, it's similar to how one wouldn't expect a deep personal connection with a customer service professional.
Some will make some profit as a niche thing (millions of users on a global scale, and if unit economics work, can make millions of $)
But it seems it will never be something really mainstream because most normal people don't care what a bot says or does.
The example I always think of is chess bots have been better at chess than humans for decades. But very few people watch stockfish tournaments. Everyone loves Magnus Carlsen though.
This is 100x for emotional support type things.