Live data from Hacker News

GPT-5 is behind schedule

wsj.com

851–860 of 1001 posts

Re: GPT-5 is behind schedule

#851

Earlier quoted context omitted.

How is synthetic data supposed to work? Broadly speaking, ML is about extracting signal from noisy data and learning the subtle patterns. If there is untapped signal in existing datasets, then learning processes should be improved. It does not follow that there should be a separate economic step where someone produces "synthetic data" from the real data, and then we treat the fake data as real data. From a scientific…

Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. I view synthetic training data like a digital twin in which it can provider further control or specified noise to understand from.

Isn’t this what Tesla does for their driving data? However it would fall apart if they didn’t have real world days to feed into it, right?

Re: GPT-5 is behind schedule

#852

Earlier quoted context omitted.

How is synthetic data supposed to work? Broadly speaking, ML is about extracting signal from noisy data and learning the subtle patterns. If there is untapped signal in existing datasets, then learning processes should be improved. It does not follow that there should be a separate economic step where someone produces "synthetic data" from the real data, and then we treat the fake data as real data. From a scientific…

Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. I view synthetic training data like a digital twin in which it can provider further control or specified noise to understand from.

What makes you assume your digital twin is actually capturing the factors that contribute to variation in the real data? This is a big issue in simulation design but for ml researchers its hand-waved off seemingly.

Re: GPT-5 is behind schedule

#853

Earlier quoted context omitted.

They’re garbage, they will always be garbage. Changing a 4 to a 5 will not make it not garbage. The whole sector is a hype bubble artificially inflating stock prices.

https://www.youtube.com/watch?v=lFc1jxLHhyM

If that’s supposed to be impressive, it really isn’t.

Re: GPT-5 is behind schedule

#854
post #753
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

They aren't blocking anything. They are just asking nicely not to be crawled. Given that AI companies haven't cared a single bit about ripping of other's peoples data I don't see why they would care now.

They block plenty and they do it crudely. I get suspicious traffic bans from reddit all the time. Trivial enough to route around by switching user agent however. Which goes to show any crawling bot writer worth their salt already routes around reddit and most other sites bs by now. I’m just the one getting the occasional headache because I use firefox and block ads and site tracking I guess.

Re: GPT-5 is behind schedule

#855
post #771

Earlier quoted context omitted.

So are you saying, some gen Z men are in a relationship, but don't know it? I do buy that, it seems to be the basis of some rom-com plots. The clueless guy that doesn't know he's being reeled in. Other factor. As the other post suggested. There are large age gaps. Women date older, men date younger. This is also long known. Does it add up to 60/30? That does seem high, but maybe with every other factor thrown in, it…

> So are you saying, some gen Z men are in a relationship, but don't know it? Or, you know, are leading women on.

Both could be happening. Guess if we are assigning some guilt, then it would depend on self awareness?

Re: GPT-5 is behind schedule

#856
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

It ultimately doesn't matter because a fairly current snapshot of all of the world's information is already housed in their data lakes. The next stage for AI training is to generate synthetic data either by other AI or by simulations to further train on as human generated content can only go so far.

[deleted]

Re: GPT-5 is behind schedule

#857
post #648

Earlier quoted context omitted.

Honestly I wish you people would stop forcing this "AI revolution" on us. It's not good. It's not useful. It's not creating value. It's not "another team member"; other team members have their own minds with their own ideas and their own opinions. Your autocomplete takes my attention away from what I want to write and replaces it with what you want me to write. We don't want it.

OP's talking about a specific use-case related to tech companies like Google. Not creative writing or research, areas in which AI is in no shape for supporting humans with it's current safety alignment.

I'm not talking about creative writing or reearch.

Re: GPT-5 is behind schedule

#858
post #741

25% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…

Using the real world- as in vision, 3d orientation, physical sensors and building training regimes that augment the language models to be multidimensional and check that perception, that is the next step. And there is very little shortage of data and experience in the actual world, as opposed to just the text internet. Can the current AI companies pivot to that? Or do you need to be worldlabs, or v2 of worldlabs?

Some can. Google owns Waymo and runs Streetview, they're collecting massive amounts of spatial data all the time. It would be harder for the MS/OpenAI centaur.

Re: GPT-5 is behind schedule

#859

Earlier quoted context omitted.

They aren't just in the hands of big corporations though. The open source, local LLM community is absolutely buzzing right now. Yes, the big companies are making the models, but enough of them are open weights that they can be fine tuned and run however you like. I think LLMs genuinely do present an opportunity to be neutral experts, or at the least neutral third parties. If they're run in completely transparent ways…

> Yes, the big companies are making the models, but enough of them are open weights that they can be fine tuned and run however you like. And how long is that going to last? This is a well known playbook at this point, we'd be better off if we didn't fall for it yet again - it's comical at this point. Sooner or later they'll lock the ecosystem down, take all the free stuff away and demand to extract the market value…

How will they do this?

You can't take the free stuff away. It's on my hard drive.

They can stop releasing them, but local models aren't going anywhere.

Re: GPT-5 is behind schedule

#860
post #479

Earlier quoted context omitted.

How do you know the answers are correct? More than once I got eloquent answer that are completely wrong.

How do you address this problem with people? More than once a real live person has told me something that was wrong,

I can draw on my past experience of interacting with the person to assign a probability to their answer being correct. Every single person in the world does this in every single human interaction they partake in, usually subconsciously.

I can't do this with an LLM because it does not have identity and may make random mistakes.

LLMs also lack the ability to say "I don't know", which my fellow humans have.

Post reply on HN