Earlier quoted context omitted.
For me the useful intuition is that LLMs haven't somehow magickally learned to implement any of the algorithms we know that we have used to make strong chess engines: alpha-beta minimax and Monte-Carlo Tree Search on the one hand, and obviously the ability to learn accurate evaluation functions by self-play. I mean we've done all this before in a task-specific fashion. It's useful to know that LLMs haven't managed to…
But it speaks in words, therefore it must be super duper extra smart!!11 /s Sarcasm aside, I think this is an easy cognitive trap to fall into. It does sometimes feel like the LLM must have some world model because it converses somewhat coherently. Examples like this failure to understand chess, or to count the number of Rs in "strawberry", seem difficult to explain if the models are intelligent. But that doesn't sto…
Why I'm still bearish on LLMs after Navier-Stokes
651–660 of 661 posts
Re: Why I'm still bearish on LLMs after Navier-Stokes
#652Earlier quoted context omitted.
No, I don't agree it's the norm. There is though a general tendency to opine with strong views on subjects posters have no expertise on. I think that's because many are software engineers (or equivalent) and they are used to being expected to "wing it" on whatever technical subject comes up. On the other hand you can always find informed comments by users who have specialist knowledge. And there's plenty of pushback…
It's not been my personal experience on this website in the last 2-3 years atleast, pre-covid perhaps. But despite that you aren't wrong and the only reason I even visit this website is because people sometimes did/do take time to reflect on things based on their experience and knowledge. And in hindsight pointing out that hn has issues wasn't even the point but I feel frustrated when everyone is readily agreeing to…
Not to disappoint you but I don't think they ever will. HN is free to join and use so people will join and use it and say whatever they want to say whether it makes sense or not. Filtering out noise is a useful skill to have especially since one can't block users or mute conversations and so on.
>> In the moment I probably thought if they are GM level and I can beat them, is this some interesting find, my disappointment honestly led me to making a rather incorrect call on this one.
Sorry, I didn't get this? What was the incorrect call you made?
Re: Why I'm still bearish on LLMs after Navier-Stokes
#653Earlier quoted context omitted.
i think you miss the part about the loop. you're still building software for the past/current world. when you have chains of agents working seamlessly, where agents are managed by themselves (like spawning a new chain to do some scope of work), and it just sort of works on a infinite game loop, your system continually keeps improving as it works. at least, that is the goal with the systems i like to build. you have t…
Your use case totally makes sense. The non-sense is the part where you think putting agents in a loop is going to yield better results indefinitely...
life of a w2
Re: Why I'm still bearish on LLMs after Navier-Stokes
#654> the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers That's a reason to be bearish about AI companies, not LLMs. But is it even true? OpenAI and Anthropic have each reported ~50 billion in revenue with ~900 billion valuations. That's a high ratio but I'm not sure if follows that the on…
I looked at the math and I think it's true. Remember revenue is just sales, not profit. These labs are shooting for > $1T valuations, which traditionally means your PROFIT is at least 1/20th or 1/30th of that (so let's say minimum 30B$/year PROFIT). These companies however are LOSING money (anthropic tries to make it sound like it's profit by deviating from accepted accounting principles) and subsidizing these models…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#655Earlier quoted context omitted.
> 98% of people think AI means what gemini tells them when they do a google search. > of that 2% who go beyond... maybe 20% of those are using AI to code. > so of that 20%, maybe 5% have... maybe 1-5% of that 5% actually went deeper How can you write any of this when you said you just started liking LLMs with Astra, a model that released two weeks ago. You're whipping out a bunch of made up stats trying to describe l…
I think he is someone who got burned by some bad software contractors in the past...
already been through one IPO where i was an early employee so i’ve never really had to work with morons.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#656Earlier quoted context omitted.
Your use case totally makes sense. The non-sense is the part where you think putting agents in a loop is going to yield better results indefinitely...
i guess me and my customers are morons then and you’re a genius when you’ve not achieved anything life of a w2
Re: Why I'm still bearish on LLMs after Navier-Stokes
#657Re: Why I'm still bearish on LLMs after Navier-Stokes
#658Re: Why I'm still bearish on LLMs after Navier-Stokes
#659This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…
As a programming lead on a hobbyist video game project, we're reaping massive rewards being in category #1. I keep the architecture and important details in check while letting frontier models go ham. As alluded, it's a video game, not a life support system, so bugs are low-impact. But even better, the defect rate is actually the lowest it's ever been. Insidious bugs baked in by years of accumulated human error are t…
But that's digression. The short of it is I've seen enough to convince me these tools may as well be magic and a whole lot of tasks that break down into "produce media content of some sort" that has a well-defined goal and definition of correctness will be permanently sped up by automation. This includes a lot of software writing. At the same time, I shared the skepticism of estimates of economic impact and irrevocably changing the larger world. I'm a lot closer to the business side of the house these days, working with customers and prospective customers to identify use cases, reference architectures, pain points, feature requests, and bring this back to the development teams to attempt using real-world experience like this to inform how we design products. It's not product management as I'm focused more often on the nitty gritty technical details, not high-level user experience or roadmaps. But it gives me a great avenue into seeing what causes organizations to actually buy and/or adopt new software products, and the rate at which they can do that.
And frankly, it isn't moving the needle much. They have the same budgets they always had, so they're not buying more, and our business is growing, but no faster or better than it grew before agentic coding became a thing. I always wonder because it seems the glowing success stories on Hacker News come in one of three varieties. It's the solo indie dev, usually targeting mobile app stores, who churns out dozens of roughly "will compile and doesn't immediately crash at runtime" apps in the time it used to take to complete one. It's the hobbyist, making software only they and maybe their immediate friends will ever use. Or it's startups, whose monetization model isn't monetization at all; it's just having something shiny to show investors in order to convince those with loose money to give enough to you personally that you can build up a nest egg whether or not your product ultimately ends up ever having a single paying customer.
In my own business, a multi-decade, mature but not hyperscale company selling overwhelmingly self-hosted enterprise open source software, I can see the impacts on output. We have the same major products with the same release cadence. Each point release averages more new features than they used to, but also more regressions. It's overall a mixed bag. Non-technical product management staff is able to contribute code. We have a ton of new internal tools that nobody uses but they're there now. On the customer side, those that hinge decisions on wanting features that didn't exist yet are benefiting from getting those. Those that already had the features they want are losing from the greater rate at which regressions get through. The net business impact seems to be things have definitely changed qualitatively, but in purely financial quantitative terms, things are about the same as they were before. More code being committed to various git forges, but same headcount, same revenue, same margins, and same market cap.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#660Earlier quoted context omitted.
Do I care if Richard Feynman was only able to do nuclear physics with the help of an abacus? Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There's joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm. Sometimes you care about the Bee, sometimes you care about the results.
This is called a benchmark. We run a calculation of Pi to evaluate a computer's performance, but we don't allow the script to download a ready-made solution. When we evaluate a runner, we don't let them use a bicycle. When we evaluate a new LLM, we don't allow it to send a request to a team of programmers, so using a chess engine for an LLM is cheating
The LLM has a process to beat chess.
Just like, if it doesn't inherently know how to multiply 13 * 17 without using Python to do it... I don't really care.
Maybe you do care. Maybe you want an LLM to be able to do work, only in its head.
But I kind of can't understand the desire for that limitation...
I mean, I do. But it seems ridiculously arbitrary. Like driving a car in 2nd gear and complaining that it gets terrible mileage and can't go fast enough. The Drive gear is literally right there.