25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist
GPT-5 is behind schedule
801–810 of 1001 posts
Re: GPT-5 is behind schedule
#80225% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist
Re: GPT-5 is behind schedule
#80325% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
IMO this is an underappreciated advantage for Google. Nobody wants to block the GoogleBot, so they can continue to scrape for AI data long after AI-specific companies get blocked. Gemini is currently embarrassingly bad given it came from the shop that: 1. invented the Transformer architecture 2. has (one of) the largest compute clusters on the planet 3. can scrape every website thanks to a long-standing whitelist
This only remains true as long as website operators think that Google Search is useful as a driver of traffic. In tech circles Google Search is already considered a flaming dumpster heap, so let's take bets on when that sentiment percolates out into the mainstream.
Re: GPT-5 is behind schedule
#80425% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
Re: GPT-5 is behind schedule
#805Earlier quoted context omitted.
In the not so distant past we already had a tool that allowed us to look up any question that came into our minds. It was super fast and always provided you with sources. It never hallucinated. It was completely free except for some advertisement. You could build a whole career out of being good at using it. It was a search engine. Young people might not remember but there was a time when Google wasn't shite but actu…
> we already had a tool that allowed us to look up any question that came into our minds … It never hallucinated. … It was a search engine. Except for all the times the search results were wrong answers. https://searchengineland.com/when-google-gets-it-wrong-direc...
At least you could check whether a source was reputable and where the bias was. With LLM's the connection between the answer and the source is completely lost. You can't even tell why it answered a certain way.
Re: GPT-5 is behind schedule
#806Earlier quoted context omitted.
A Call to expertise is actually a fallacy. This is because experts can be wrong. The scientific method relies on evidence and reproducible results, not authority alone. Edited to add a reference: see under Appeal to authority. https://writingcenter.unc.edu/tips-and-tools/fallacies/
The fact is that in science, facts are only definitions and everything else is a theory which by definition is never 100% true.
Which definition of theory includes that it can never be 100% true? It can't be proven to be true, but surely it could be true without anyone knowing about it.
Re: GPT-5 is behind schedule
#807Earlier quoted context omitted.
Same here. The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It reminds me of being a kid and asking my grandpa a million questions, like how light bulbs worked, or what was inside his radio, or how do we have day and night. And before anyone talks about accuracy or hallucinations, these conversations usually are treated as starting off poi…
> The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It is dangerous to assume that LLMs are experts on any topic. With or without quotes. You are getting a super fast journalist intern with a huge memory but inability to reason critically, lacking understanding about anything and huge unreliability when it comes to answering questions (you…
I treat LLM answers about the same way I treat wikipedia articles. If it's critical I get it right, I go to the wiki sources referenced. Recent models have gotten good at 'showing their sources', which is helpful.
Re: GPT-5 is behind schedule
#80825% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
They aren't blocking anything. They are just asking nicely not to be crawled. Given that AI companies haven't cared a single bit about ripping of other's peoples data I don't see why they would care now.
Re: GPT-5 is behind schedule
#80925% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
To your point, I have wondered whatever became of that massive initiative from Google to scan books, and whether that might be looked at as a potential training source, giving that Google has run into legal limitations on other forms of usage.
Re: GPT-5 is behind schedule
#810Earlier quoted context omitted.
At this point it’s quite likely that they could pivot and just be the chatgpt company. I’ve found chatgpt-4o with web search and plugins to be more useful than o1 for most tasks. It’s possible we’re nearing the end of the LLM race, but I doubt that’s the end of the AI story this decade, or OpenAI.
Ya I think they probably will, but "the chatgpt company" is not worth 157B. It might not even be worth 1B.
The cash burn and claim of imminent agi is where the valuation trouble could be.