Live data from Hacker News

A bear case: My predictions regarding AI progress

lesswrong.com

81–90 of 220 posts

Re: A bear case: My predictions regarding AI progress

#81
post #22

> LLMs still seem as terrible at this as they'd been in the GPT-3.5 age. Software agents break down once the codebase becomes complex enough, game-playing agents get stuck in loops out of which they break out only by accident, etc. This has been my observation. I got into Github Copilot as early as it launched back when GPT-3 was the model. By that time (late 2021) copilot can already write tests for my Rust function…

While I don't disagree with that observation, it falls into the "well, duh!"-category for me. The models are build with no mechanism for long term memory and thus suck at tasks that require long term memory. There is nothing surprising here. There was never any expectation that LLMs magically develop long term memory, as that's impossible given the architecture. They predict the next word and once the old text moves out of the context window, it's gone. The models neither learn as they work nor can they remember the past.

It's not even like humans are all that different here. Strip a human of their tools (pen&paper, keyboard, monitor, etc.) and have them try solving problems with nothing but the power of their brain and they'll struggle a hell of a lot too, since our memory ain't exactly perfect either. We don't have perfect recall, we look things up when we need to, a large part of our "memory" is out there in the world around us, not in our head.

The open question is how to move forward. But calling AI progress a dead end before we even started exploring long term memory, tool use and on-the-fly learning is a tad little premature. It's like calling quits on the development of the car before you put the wheels on.

Re: A bear case: My predictions regarding AI progress

#82
post #29

Let's imagine that we all had a trillion dollars. Then we would all sit around and go "well dang, we have everything, what should we do?". I think you'll find that just about everyone would agree, "we oughta see how far that LLM thing can go". We could be in nuclear fallout shelters for decades, and I think you'll still see us trying to push the LLM thing underground, through duress. We dream of this, so the bear cas…

Wdym all of us? I certainly would find much better usages for the money. What about reforming democracy? Use the corrupt system to buy the votes, then abolish all laws allowing these kind of donations that allow buying votes. I'll litigate the hell out of all the oligarchs now that they can't out pay justice. This would pay off more than a moon shot. I would give a bit of money for the moon shot, why not, but not all…

"So, after Rome's all yours you just give it back to the people? Tell me why."

Re: A bear case: My predictions regarding AI progress

#83
post #51

This seems to be ignoring the major force driving AI right now - hardware improvements. We've barely seen a new hardware generation since ChatGPT was released to the market, we'd certainly expect it to plateau fairly quickly on fixed hardware. My personal experience of AI models is going to be a series of step changes every time the VRAM on my graphics card doubles. Big companies are probably going to see something s…

> Although I would bet money on there being a few years left and AGI is achieved.

Yeah? I'll take you up on that offer. $100AUD AGI won't happen this decade.

Re: A bear case: My predictions regarding AI progress

#84
I think the author provides an interesting perspective to the AI hype, however, I think he is really downplaying the effectiveness of what you can do with the current models we have.

If you've been using LLMs effectively to build agents or AI-driven workflows you understand the true power of what these models can do. So in some ways the author is being a little selective with his confirmation bias.

I promise you that if you do your due diligence in exploring the horizon of what LLMs can do you will understand what I'm saying. If ya'll want a more detailed post I can get into the AI systems I have been building. Don't sleep on AI.

Re: A bear case: My predictions regarding AI progress

#85
post #32

The thing I can't wrap my head around is that I work on extremely complex AI agents every day and I know how far they are from actually replacing anyone. But then I step away from my work and I'm constantly bombarded with “agents will replace us”. I wasted a few days trying to incorporate aider and other tools into my workflow. I had a simple screen I was working on for configuring an AI Agent. I gave screenshots of…

There are some fields though where they can replace humans in significant capacity. Software development is probably one of the least likely for anything more than entry level, but A LOT of engineering has a very very real existential threat. Think about designing buildings. You basically just need to know a lot of rules / tables and how things interact to know what's possible and the best practices. A purpose built…

Most engineering fields are de jure professional, which means they can and probably will enforce limitations on the use of GenAI or its successor tech before giving up that kind of job security. Same goes for the legal profession.

Software development does not have that kind of protection.

Re: A bear case: My predictions regarding AI progress

#86
post #60

Earlier quoted context omitted.

> I knew Uber, Netflix, Spotify were revolutionary the first time I used them. Maybe re-tune your revolution sensor. None of those are revolutionary companies. Profitable and well executed, sure, but those turn up all the time. Uber's entire business model was running over the legal system so quickly that taxi licenses didn't have time to catch up. Other than that it was a pretty obvious idea. It is a taxi service. T…

They were revolutionary as product genres, not necessary individual companies. Ordering a cab without making a phone call was revolutionary. Netflix at least with its initial promise of having all the world's movies and TV was revolutionary, but it didn't live up to that. Spotify because of how cheap and easy it was to have access to all the music, this was the era when people were paying 99c per song on iTunes. I've…

"Do something existing with a different mechanism" is innovative, but not revolutionary, and certainly not a new "product genre". My parents used to order pizza by phone calls, then a website, then an app. It's the same thing. (The friction is a little bit less, but maybe forcing another human to bring food to you because you're feeling lazy should have a little friction. And as a side effect, we all stopped being as comfortable talking to real people on phone calls!)

Napster came before Spotify.

Re: A bear case: My predictions regarding AI progress

#87
post #67
post #60

Earlier quoted context omitted.

> I knew Uber, Netflix, Spotify were revolutionary the first time I used them. Maybe re-tune your revolution sensor. None of those are revolutionary companies. Profitable and well executed, sure, but those turn up all the time. Uber's entire business model was running over the legal system so quickly that taxi licenses didn't have time to catch up. Other than that it was a pretty obvious idea. It is a taxi service. T…

> None of those are revolutionary companies. Not only Uber/Grab (or delivery app) were revolutionary, they are still revolutionary. I could live without LLMs and my life will be slightly impacted when coding. If delivery apps are not available, my life is severely degraded. The other day I was sick. I got medicine and dinner with Grab. Delivered to the condo lobby which is as far as I can get. That is revolutionary.

Is it revolutionary to order from a screen rather than calling a restaurant for delivery? I don’t think so.

Re: A bear case: My predictions regarding AI progress

#88
post #67
post #60

Earlier quoted context omitted.

> I knew Uber, Netflix, Spotify were revolutionary the first time I used them. Maybe re-tune your revolution sensor. None of those are revolutionary companies. Profitable and well executed, sure, but those turn up all the time. Uber's entire business model was running over the legal system so quickly that taxi licenses didn't have time to catch up. Other than that it was a pretty obvious idea. It is a taxi service. T…

> None of those are revolutionary companies. Not only Uber/Grab (or delivery app) were revolutionary, they are still revolutionary. I could live without LLMs and my life will be slightly impacted when coding. If delivery apps are not available, my life is severely degraded. The other day I was sick. I got medicine and dinner with Grab. Delivered to the condo lobby which is as far as I can get. That is revolutionary.

Were you not able to order food before Uber/Grab?

Re: A bear case: My predictions regarding AI progress

#89
post #88
post #67

Earlier quoted context omitted.

> None of those are revolutionary companies. Not only Uber/Grab (or delivery app) were revolutionary, they are still revolutionary. I could live without LLMs and my life will be slightly impacted when coding. If delivery apps are not available, my life is severely degraded. The other day I was sick. I got medicine and dinner with Grab. Delivered to the condo lobby which is as far as I can get. That is revolutionary.

Were you not able to order food before Uber/Grab?

I am not in the US and yes, it is not a thing (though there was a pizza place that had phone order, but that's rather an exception).

Re: A bear case: My predictions regarding AI progress

#90
post #25

Earlier quoted context omitted.

Nope. I try the latest models as they come and I have a self-made custom setup (as in a custom lua plugin) in Neovim. What I am not, is selling AI or AI-driven solutions.

Do you mean that you have successfully managed to get the same experience in cursor but in neovim? I have been looking for something like that to move back to my neovim setup instead of using cursor. Any hints would be greatly appreciated!

Start with Avante or CopilotChat. Create your own Lua config/plugin (easy with Claude 3.5 ;) ) and then use their chat window to run copilot/models. Most of my custom config was built with Claude 3.5 and some trial/error/success.
Post reply on HN