Live data from Hacker News

GPT-5 is behind schedule

wsj.com

811–820 of 1001 posts

Re: GPT-5 is behind schedule

#811
post #796
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist

There are two to distinguish: "Googlebot" and "Google-Extended".

Re: GPT-5 is behind schedule

#812

Earlier quoted context omitted.

> a revolution in knowledge-worker productivity. That's a nice euphemism for "imminent mass layoffs and a race to the bottom"...

No, the job market will adapt, just like it did during the industrial and information revolutions, and life will be better.

It will be better for those who already have it good. How it will affect those who don't is the real question here.

Re: GPT-5 is behind schedule

#813
post #796
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist

For OpenAI, they could lean on their relationship with Microsoft for Bing crawler access

Websites won’t be blocking the search engine crawlers until they stop sending back traffic, even if they’re sending back less and less traffic

Re: GPT-5 is behind schedule

#814
post #803
post #796

Earlier quoted context omitted.

IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist

> Nobody wants to block the GoogleBot This only remains true as long as website operators think that Google Search is useful as a driver of traffic. In tech circles Google Search is already considered a flaming dumpster heap, so let's take bets on when that sentiment percolates out into the mainstream.

If it reaches the point where google is no longer a useful driver of traffic then there's probably little point in having a website at all any more.

Re: GPT-5 is behind schedule

#815
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

It ultimately doesn't matter because a fairly current snapshot of all of the world's information is already housed in their data lakes. The next stage for AI training is to generate synthetic data either by other AI or by simulations to further train on as human generated content can only go so far.

Re: GPT-5 is behind schedule

#816
post #749

Earlier quoted context omitted.

> we already had a tool that allowed us to look up any question that came into our minds … It never hallucinated. … It was a search engine. Except for all the times the search results were wrong answers. https://searchengineland.com/when-google-gets-it-wrong-direc...

Being biased is not the same as hallucinating. LLMs have both problems. At least you could check whether a source was reputable and where the bias was. With LLM's the connection between the answer and the source is completely lost. You can't even tell why it answered a certain way.

> Being biased is not the same as hallucinating. LLMs have both problems.

I didn't deny either of those things, I said that search engines also hallucinate — my actual link gave several examples, including "King of the United States" -> "Barack Obama".

Just because it showed the link to breitbart doesn't mean it was not hallucinating.

> At least you could check whether a source was reputable and where the bias was.

The former does not imply the latter. You could tell where a search engine got an answer from, but not which answers were hidden — an argument that I saw some on the American right make to criticise Google for failing to show their version of events.

> With LLM's the connection between the answer and the source is completely lost. You can't even tell why it answered a certain way.

Also not so. The free version of ChatGPT supports search directly, so it allows you to have references.

Re: GPT-5 is behind schedule

#817

Earlier quoted context omitted.

> The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It is dangerous to assume that LLMs are experts on any topic. With or without quotes. You are getting a super fast journalist intern with a huge memory but inability to reason critically, lacking understanding about anything and huge unreliability when it comes to answering questions (you…

I actually find LLMs lacking true expertise to be a feature, not a bug. Most of the time I'm starting from a place of no knowledge on a topic that's novel to me, I ask some questions, it replies with summaries, keywords, names of things, basic concepts. I enter with the assumption that it's really no different than googling phrases and sifting through results (except I don't know what phrases I'm supposed to be googl…

You forget that it makes stuff up and you won't know it until you google it. When googling, fake stuff stands out because truth is consistent.

Querying multiple llms at the same time and being able to compare results is a much better comparison to googling but no one does this.

As I said, you are talking to a super confident journalist intern who can give you answers but you won't know if it is true or partially true until you consult with a human source of knowledge.

It's not even similar to asking the old guys at the Home Depot because they can tell you if they are unsure they have a good answer for you. An LLM won't. Old guys won't hallucinate facts the way an LLM will

It is really is the 21st century Searle's epistemological Chinese room nightmare edition. Grammar checks out but whatever is spit out doesn't necessarily bear any resemblance to reality

Re: GPT-5 is behind schedule

#818
post #784

Earlier quoted context omitted.

It’s almost as if human intelligence doesn’t involve performing repeated matrix multiplications over a mathematically transformed copy of the internet. ;-)

It makes you think there must be more efficient algorithms out there.

Maybe the problem isn't the algorithm but the hardware. Numerically simulating the thermal flow in a lightbulb or CFD of a Stone flying through air is pretty hard, but the physical thing isn't that complex to do. We're trying to simulate the function of a brain which is basically an analog thing using a digital computer. Of course that can be harder than running the brain itself.

Re: GPT-5 is behind schedule

#819

I just want to say something random. ChatGPT cited and English newspaper called 'The Sun' in the answer to a query I made this moring. I found that amusing and figured then that we might already be beyond peak AI! If AI is trained on sources that include The Sun, it's going to end up being garbage and/or constantly needing to be fact checked.

It is, which shouldn't be a surprise to anyone.

Re: GPT-5 is behind schedule

#820

"Orion’s problems signaled to some at OpenAI that the more-is-more strategy, which had driven much of its earlier success, was running out of steam." So LLMs finally hit the wall. For a long time, more data, bigger models, and more compute to drive them worked. But that's apparently not enough any more. Now someone has to have a new idea. There's plenty of money available if someone has one. The current level of LLM…

The new idea is already here and it's reasoning / chain of thought. Anecdotally Claude is pretty good at knowing the bounds of its knowledge.

Not even close. I’m a programmer but also a guitarist. I love asking it to tab out songs for me or asking it how many bars are in the intro of a song. It convincingly gives an answer that is always way off the mark.
Post reply on HN