Live data from Hacker News

GPT-5 is behind schedule

wsj.com

831–840 of 1001 posts

Re: GPT-5 is behind schedule

#831
post #803

Earlier quoted context omitted.

> Nobody wants to block the GoogleBot This only remains true as long as website operators think that Google Search is useful as a driver of traffic. In tech circles Google Search is already considered a flaming dumpster heap, so let's take bets on when that sentiment percolates out into the mainstream.

If it reaches the point where google is no longer a useful driver of traffic then there's probably little point in having a website at all any more.

Strange take ... I seem to remember websites having a lot of point before google.

Re: GPT-5 is behind schedule

#832
post #35

Earlier quoted context omitted.

That so weird, it’s seems like everybody here prefers Claude. I’ve been using Claude and openai in copilot and I find even 4o seems to understand the problem better. O1 definitely seems to get it right more for me.

I try to sprinkle 'for us/me' everywhere as much as I can; we work on LoB/ERP apps mostly. These are small frontends to massive multi million line backends. We carved a niche by providing the frontends on these backends live at the client office by a business consultant of ours: they simply solve UX issues for the client on top of large ERP by using our tool and prompting. Everything looks modern, fresh and nice; unl…

> It's fast and no frontend people are needed for it

I guess if you don’t need to maintain it, just an ever growing blob of complexity that will be reinvented into new blobs every time when the old one becomes too immobile :)

Re: GPT-5 is behind schedule

#833
post #823

Earlier quoted context omitted.

If you think of human neurons they seem to basically take inputs from bunch of other neurons, possibly modified by chemical levels and send out a signal when they get enough. It seems like something that could be functionally simulated in software by some fairly basic adding up inputs type stuff rather than needing the details of all the chemistry.

Isn’t that exactly what we’re currently doing? The problem is that doing this few billion times for every token seems to be harder than just powering some actual neurons with sugar.

I think the algorithm is pretty different though I'm not expert on the stuff. I don't think the brain processes look like matrix multiplication.

Re: GPT-5 is behind schedule

#834
post #831

Earlier quoted context omitted.

If it reaches the point where google is no longer a useful driver of traffic then there's probably little point in having a website at all any more.

Strange take ... I seem to remember websites having a lot of point before google.

They had a point back then because no alternatives existed.

How many websites back then would be youtube channels, podcasts or social media accounts if they had existed back then?

Nowadays most sites survive via traffic from google, if it goes away then most of those sites go away as well.

Re: GPT-5 is behind schedule

#835
post #753

Earlier quoted context omitted.

They aren't blocking anything. They are just asking nicely not to be crawled. Given that AI companies haven't cared a single bit about ripping of other's peoples data I don't see why they would care now.

A number of sites have started outright blocking any traffic that looks remotely suspicious. This has made browsing with a vpn a bit of a pain.

This has been ever increasing for years now. Bots, attacks, scrapers, AI, all these things seem to be the majority of traffic on most sites.

Re: GPT-5 is behind schedule

#836
post #829
post #585

Earlier quoted context omitted.

Tesla is still valued high, despite FSD did not came, despite being promised. So OpenAI would get away with delivering ChatGPT5, if it is better than the competition.

Tesla is profitable and they have a big technological moat. OpenAI is in a very competitive industry and they burn ~5B a year.

I believe the car industry is somewhat competive as well and they needed allmost 10 years to become profitable.

Re: GPT-5 is behind schedule

#837
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

Doing basic copyright analyses on model outputs is all that is needed. Check if the output contains copyright, block it if it does. Transformers aren't zettabyte sized archives with a smart searching algo, running around the web stuffing everything they can into their datacenter sized storage. They are typically a few dozen GB in size, if that. They don't copy data, they move vectors in a high dimensional space based…

It's not even close to that simple. Nobody is really questioning if the data contains the copyrighted information, we know that to be true in enough cases to bankrupt open ai, the question is what analogy should the courts be using as a basis to determine if it's infringement.

It read many works but can't duplicate them exactly sounds a lot like what I've done, to be honest. I can give you a few memorable lines to a few songs but only really can come close to reciting my favorites completely. The LLMs are similar but their favorites are the favorites of the training data. A line in a pop song mentioned a billion times is likely reproducible, the lyrics to the next track on the album, not so much.

IMO, any infringement that might have happened would be acquiring data in the first place but copy protection cares more about illegal reproduction than illegal acquisition.

Re: GPT-5 is behind schedule

#838
post #686

Earlier quoted context omitted.

People who are experts (PhD and 20 years of experience) often have very dumb opinions in their field of expertise. Experts make amateur mistakes too. Look at the books written by expert economists, expert psychologists, expert historians, expert philosophers, expert software engineers. Most books are not worth the paper they're written on, despite the authors being experts with decades of experience in their respecti…

Right, but, then what? If you throw away all of the books from experts, what do you do, go out in your backyard and start running experiments to re-create all of science? Or start googling? What, some random person on the internet is going to be a better 'expert' than someone that wrote a book? Books might not be great, but they are at least some minimum bar to reach. You had to do some study and analysis. Seems like…

Many terrific books have been published in the past 500 years. The median book is not worth your time, however, and neither is the top 10%. You cannot possibly read everything so you have to be very selective or you will read only dreck. This is the opposite of being anti-science or anti-education.

Re: GPT-5 is behind schedule

#839
post #824

Earlier quoted context omitted.

>Bill Gross correctly calls this phase of AI shoplifting. I call it the Napster-of-Everything (because I am old). I am also betting that the courts won't buy the "fair use" interpretation of scraping, given the revenues AI companies generate. That means a potential stalling of new models until some mechanism is worked out to pay knowledge creators. To your point, I have wondered whatever became of that massive initia…

> To your point, I have wondered whatever became of that massive initiative from Google to scan books, and whether that might be looked at as a potential training source, giving that Google has run into legal limitations on other forms of usage. Still around, doing fine: https://en.wikipedia.org/wiki/Google_Books and https://books.google.com/intl/en/googlebooks/about/index.htm... Given the timing, I suspect it was st…

I don't know what you mean by timing (relative to what?) or "simple indexing" (they scanned the complete contents of books), but I am, and was already aware, of the wiki article and the role of recaptcha.

Maybe I wasn't clear, but I was interested in the consequences of the legal stuff. It's not clear from the wiki article what any of this means with respect to the suitability of scans for AI training.

Re: GPT-5 is behind schedule

#840
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

It ultimately doesn't matter because a fairly current snapshot of all of the world's information is already housed in their data lakes. The next stage for AI training is to generate synthetic data either by other AI or by simulations to further train on as human generated content can only go so far.

How is synthetic data supposed to work? Broadly speaking, ML is about extracting signal from noisy data and learning the subtle patterns.

If there is untapped signal in existing datasets, then learning processes should be improved. It does not follow that there should be a separate economic step where someone produces "synthetic data" from the real data, and then we treat the fake data as real data. From a scientific perspective, that last part sounds really bad.

Creating derivative data from real data sounds, for the purpose of machine learning, like a scam by the data broker industry. What is the theory behind it, if not fleecing unsophisticated "AI" companies? Is it just myopia, Goodhart's Law applied to LLM scaling curves? Some MBA took the "data is the new oil" comment a little too seriously and inferred that data is as fungible as refined petroleum?

Post reply on HN