Live data from Hacker News

GPT-5 is behind schedule

wsj.com

801–810 of 1001 posts

Re: GPT-5 is behind schedule

#801
post #796
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist

Wonder if OpenAI is considering building a search engine for this reason... Imagine if we get a functional search engine again from some company just trying to feeding their model generation...

Re: GPT-5 is behind schedule

#802
post #796
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist

The new Gemini Experimental models are the best general purpose models out right now. I have been comparing with o1 Pro and I prefer Gemini Experimental 1206 due to its context, speed, and accuracy. Google came out with a lot of new stuff last week if you havent been following. They seem to have the best models across the board, including image and video.

Re: GPT-5 is behind schedule

#803
post #796
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist

> Nobody wants to block the GoogleBot

This only remains true as long as website operators think that Google Search is useful as a driver of traffic. In tech circles Google Search is already considered a flaming dumpster heap, so let's take bets on when that sentiment percolates out into the mainstream.

Re: GPT-5 is behind schedule

#804
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

With current state of legal, a real challenge can happen only around 10 years from now. By then AI players will gather immense power over the law.

Re: GPT-5 is behind schedule

#805
post #749

Earlier quoted context omitted.

In the not so distant past we already had a tool that allowed us to look up any question that came into our minds. It was super fast and always provided you with sources. It never hallucinated. It was completely free except for some advertisement. You could build a whole career out of being good at using it. It was a search engine. Young people might not remember but there was a time when Google wasn't shite but actu…

> we already had a tool that allowed us to look up any question that came into our minds … It never hallucinated. … It was a search engine. Except for all the times the search results were wrong answers. https://searchengineland.com/when-google-gets-it-wrong-direc...

Being biased is not the same as hallucinating. LLMs have both problems.

At least you could check whether a source was reputable and where the bias was. With LLM's the connection between the answer and the source is completely lost. You can't even tell why it answered a certain way.

Re: GPT-5 is behind schedule

#806
post #745

Earlier quoted context omitted.

A Call to expertise is actually a fallacy. This is because experts can be wrong. The scientific method relies on evidence and reproducible results, not authority alone. Edited to add a reference: see under Appeal to authority. https://writingcenter.unc.edu/tips-and-tools/fallacies/

The fact is that in science, facts are only definitions and everything else is a theory which by definition is never 100% true.

> everything else is a theory which by definition is never 100% true.

Which definition of theory includes that it can never be 100% true? It can't be proven to be true, but surely it could be true without anyone knowing about it.

Re: GPT-5 is behind schedule

#807
post #457

Earlier quoted context omitted.

Same here. The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It reminds me of being a kid and asking my grandpa a million questions, like how light bulbs worked, or what was inside his radio, or how do we have day and night. And before anyone talks about accuracy or hallucinations, these conversations usually are treated as starting off poi…

> The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It is dangerous to assume that LLMs are experts on any topic. With or without quotes. You are getting a super fast journalist intern with a huge memory but inability to reason critically, lacking understanding about anything and huge unreliability when it comes to answering questions (you…

Shrug

I treat LLM answers about the same way I treat wikipedia articles. If it's critical I get it right, I go to the wiki sources referenced. Recent models have gotten good at 'showing their sources', which is helpful.

Re: GPT-5 is behind schedule

#808
post #753
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

They aren't blocking anything. They are just asking nicely not to be crawled. Given that AI companies haven't cared a single bit about ripping of other's peoples data I don't see why they would care now.

Wouldn't it be somewhat trivial to set up honeypots?

Re: GPT-5 is behind schedule

#809
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

>Bill Gross correctly calls this phase of AI shoplifting. I call it the Napster-of-Everything (because I am old). I am also betting that the courts won't buy the "fair use" interpretation of scraping, given the revenues AI companies generate. That means a potential stalling of new models until some mechanism is worked out to pay knowledge creators.

To your point, I have wondered whatever became of that massive initiative from Google to scan books, and whether that might be looked at as a potential training source, giving that Google has run into legal limitations on other forms of usage.

Re: GPT-5 is behind schedule

#810
post #485
post #461

Earlier quoted context omitted.

At this point it’s quite likely that they could pivot and just be the chatgpt company. I’ve found chatgpt-4o with web search and plugins to be more useful than o1 for most tasks. It’s possible we’re nearing the end of the LLM race, but I doubt that’s the end of the AI story this decade, or OpenAI.

Ya I think they probably will, but "the chatgpt company" is not worth 157B. It might not even be worth 1B.

Id be hard pressed to come up with a valuation under 30B based on the publicly known finances. OpenAI is certainly crushing the metrics of other highly valued startups like snowflake and databricks.

The cash burn and claim of imminent agi is where the valuation trouble could be.

Post reply on HN