Earlier quoted context omitted.
LLMs aren't good at being search engines, they're good at understanding things. Put an LLM on top of a search engine, and that's the appropriate tool for this use case. I guess the problem with LLMs is that they're too usable for their own good, so people don't realizing that they can't perfectly know all the trivia in the world, exactly the same as any human.
Ironically though an LLM powered search engine (some word about being perplexed) is becoming way better than the undisputed king of traditional search engines (something oogle)
Recent AI model progress feels mostly like bullshit
101–110 of 478 posts
Re: Recent AI model progress feels mostly like bullshit
#102Earlier quoted context omitted.
This seems like a probable end state, but we're going to have to stop calling LLMs "artificial intelligence" in order to get there.
Why not? Objectively speaking LLMs are artificial intelligent. Just because it's not human level intelligence doesn't mean it's not intelligent.
The fact is, the phrase "artificial intelligence" is a memetic hazard: it immediately positions the subject of conversation as "default capable", and then forces the conversation into trying to describe what it can't do, which is rarely a useful way to approach it.
Whereas with LLMs (and chess engines and every other tech advancement) it would be more useful to start with what the tech _can_ do and go from there.
Re: Recent AI model progress feels mostly like bullshit
#103Earlier quoted context omitted.
This is less an LLM thing than an information retrieval question. If you choose a model and tell it to “Search,” you find citation based analysis that discusses that he indeed had problems with alcohol. I do find it interesting it quibbles whether he was an alcoholic or not - it seems pretty clear from the rest that he was - but regardless. This is indicative of something crucial when placing LLMs into a toolkit. The…
Any information found in a web search about Newman will be available in the training set (more or less). It's almost certainly a problem of alignment / "safety" causing this issue.
Re: Recent AI model progress feels mostly like bullshit
#104Earlier quoted context omitted.
I asked GPT-4.5 and it searched the web and immediately gave me a "yes" with paragraphs of sources cited.
Truth is a probability game. Just keep trying until you arrive.
Re: Recent AI model progress feels mostly like bullshit
#105The disconnect between improved benchmark results and lack of improvement on real world tasks doesn't have to imply cheating - it's just a reflection of the nature of LLMs, which at the end of the day are just prediction systems - these are language models, not cognitive architectures built for generality. Of course, if you train an LLM heavily on narrow benchmark domains then its prediction performance will improve…
Re: Recent AI model progress feels mostly like bullshit
#106Earlier quoted context omitted.
LLMs aren't good at being search engines, they're good at understanding things. Put an LLM on top of a search engine, and that's the appropriate tool for this use case. I guess the problem with LLMs is that they're too usable for their own good, so people don't realizing that they can't perfectly know all the trivia in the world, exactly the same as any human.
> LLMs aren't good at being search engines, they're good at understanding things. LLMs are literally fundamentally incapable of understanding things. They are stochastic parrots and you've been fooled.
Re: Recent AI model progress feels mostly like bullshit
#107It’s not even approaching the asymptotic line of promises made at any achievable rate for the amount of cash being thrown at it. Where’s the business model? Suck investors dry at the start of a financial collapse? Yeah that’s going to end well…
> where’s the business model? For who? Nvidia sell GPUs, OpenAI and co sell proprietary models and API access, and the startups resell GPT and Claude with custom prompts. Each one is hoping that the layer above has a breakthrough that makes their current spend viable. If they do, then you don’t want to be left behind, because _everything_ changes. It probably won’t, but it might. That’s the business model
Re: Recent AI model progress feels mostly like bullshit
#108Earlier quoted context omitted.
LLMs aren't good at being search engines, they're good at understanding things. Put an LLM on top of a search engine, and that's the appropriate tool for this use case. I guess the problem with LLMs is that they're too usable for their own good, so people don't realizing that they can't perfectly know all the trivia in the world, exactly the same as any human.
> LLMs aren't good at being search engines, they're good at understanding things. LLMs are literally fundamentally incapable of understanding things. They are stochastic parrots and you've been fooled.
Take two 4K frames of a falling vase, ask a model to predict the next token... I mean the following images. Your model now needs include some approximations of physics - and the ability to apply it correctly - to produce a realistic outcome. I'm not aware of any model capable of doing that, but that's what it would mean to predict the unseen with high enough fidelity.
Re: Recent AI model progress feels mostly like bullshit
#109This is a bit of a meta-comment, but reading through the responses to a post like this is really interesting because it demonstrates how our collective response to this stuff is (a) wildly divergent and (b) entirely anecdote-driven. I have my own opinions, but I can't really say that they're not also based on anecdotes and personal decision-making heuristics. But some of us are going to end up right and some of us ar…
Re: Recent AI model progress feels mostly like bullshit
#110Earlier quoted context omitted.
> LLMs aren't good at being search engines, they're good at understanding things. LLMs are literally fundamentally incapable of understanding things. They are stochastic parrots and you've been fooled.
What does the word "understand" mean to you?
Rather than give you a technical answer - if I ever feel like an LLM can recognize its limitations rather than make something up, I would say it understands. In my experience LLMs are just algorithmic bullshitters. I would consider a function that just returns "I do not understand" to be an improvement, since most of the time I get confidently incorrect answers instead.
Yes, I read Anthropic's paper from a few days ago. I remain unimpressed until talking to an LLM isn't a profoundly frustrating experience.