Earlier quoted context omitted.
Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. I view synthetic training data like a digital twin in which it can provider further control or specified noise to understand from.
What makes you assume your digital twin is actually capturing the factors that contribute to variation in the real data? This is a big issue in simulation design but for ml researchers its hand-waved off seemingly.
GPT-5 is behind schedule
901–910 of 1001 posts
Re: GPT-5 is behind schedule
#902Earlier quoted context omitted.
How is synthetic data supposed to work? Broadly speaking, ML is about extracting signal from noisy data and learning the subtle patterns. If there is untapped signal in existing datasets, then learning processes should be improved. It does not follow that there should be a separate economic step where someone produces "synthetic data" from the real data, and then we treat the fake data as real data. From a scientific…
Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. I view synthetic training data like a digital twin in which it can provider further control or specified noise to understand from.
You're suggesting the new, untested models in a new, untested technological field are sufficient for deployment in real world applications even with a lack of real world data to supplement them. That's magical thinking given what we've experienced in every other field of engineering (and finance for that matter).
Why is AI/ML any different? Because highly anthropomorphized words like "learning" and "intelligence" are in the name? These models are some of the most complex machines humanity has ever produced. Replace "learning" and "intelligence" with "calibrated probability calculators". Then detail the sheer complexity of the calibrations needed, and tell me with a straight face that simulations are good enough.
Re: GPT-5 is behind schedule
#90325% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
> I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. This is where I'm at. I write content when I run into problems that I don't see solved anywhere else, so my sites host novel…
I work in a pretty niche field and feel the same way. I don't mind sharing my writing with individuals (even if they don't directly cite me) because then they see my name and know who came up with it, so I still get some credit. You could call this "clout farming" or something derogatory, but this is how a lot of experts genuinely get work...by being known as "the guy who gave us that great tip on a blog once".
With AI snooping around, I feel like becoming one of those old mathematicians that would hold back publicizing new results to keep them all for themselves. That doesn't seem selfish to me, humans have a right to protect ourselves and survive and maintain the value of our expertise when OpenAI isn't offering any money.
I honestly think we should just be done with writing content online now, before it's too late. I've thought a lot about it lately and I'm leaning more towards that option.
Re: GPT-5 is behind schedule
#904Earlier quoted context omitted.
Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. I view synthetic training data like a digital twin in which it can provider further control or specified noise to understand from.
> Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. No, just as I wouldn't trust a surgeon who studied medicine by playing Operation. A gross approximation is not a substitute for real life.
Re: GPT-5 is behind schedule
#905Earlier quoted context omitted.
> Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. No, just as I wouldn't trust a surgeon who studied medicine by playing Operation. A gross approximation is not a substitute for real life.
What about a doctor who used a mix of training both on live patients as well as cadavers and models?
Similarly, when this doctor sees something new, will they just write it off as something they've seen before and confidently work from that assumption?
Re: GPT-5 is behind schedule
#906Earlier quoted context omitted.
Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. I view synthetic training data like a digital twin in which it can provider further control or specified noise to understand from.
What makes you assume your digital twin is actually capturing the factors that contribute to variation in the real data? This is a big issue in simulation design but for ml researchers its hand-waved off seemingly.
https://www.forbes.com/sites/carolynschwaar/2024/12/09/schae...
Re: GPT-5 is behind schedule
#90725% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
See https://fairuse.stanford.edu/overview/fair-use/four-factors/
I think in particular it fails the "Amount and substantiality of the portion taken" and "Effect of the use on the potential market" extremely egregiously.
Re: GPT-5 is behind schedule
#908Earlier quoted context omitted.
You can turn to actual experts, e.g. YouTube or books. But yes, I have recently had the misfortune of working with a personal trainer who was using ChatGPT to come up with training programs, and it felt confusing and like I was wasting time and money.
When I'm looking for actual experts, the first thing that comes to my mind is definitely YouTube!! And least when it's about YouTube specific topics, like where the like button and the subscribe button is. They will tell me. Every. Single. F*cking. 5. Minute. Clip. Again. And. Again. Not soooo much for anything actually important or interesting, though.... ;) PS: Also which of the always same ~5 shady companies their…
>They will tell me. Every. Single. F*cking. 5. Minute. Clip. Again. And. Again.
Do you know why you got that video. Because people liked and subscribed to them and the 'experts' with the best information in the universe are hidden 5000 videos below with 10 views.
And this is 100% Googles fault for the algorithms they created that force these behaviors on anyone that wants to use their platform and have visibility.
Lastly, if you can't find anything interesting or important on YT, this points at a failure of your own. While there is an ocean of crap, there is more than enough amazing content out there.
Re: GPT-5 is behind schedule
#909Earlier quoted context omitted.
> Would you trust a ML self-driving algorithm trained on a "digital twin" of a city? I would. No, just as I wouldn't trust a surgeon who studied medicine by playing Operation. A gross approximation is not a substitute for real life.
What about a doctor who used a mix of training both on live patients as well as cadavers and models?
Re: GPT-5 is behind schedule
#910I'm sure the debate over the definition of AGI is important and will continue for a while, but... I can't care about it anymore. Between Perplexity searching and summarizing, Claude explaining, and qwen (and other tools) coding, I'm already as happy as can be with whatever you want to call this level of intelligence. Just today I used a completely local AI research tool, based on Ollama. It worked great. Maybe it won…
The definition of agi is a linguistic problem but people confuse it for a philosophical problem. Think about it. The term is basically just a classification and what features and qualities fit the classification is an arbitrary and linguistic choice. The debate stems from a delusion and failure to realize that people are simply picking and choosing different fringe features on what qualifies as agi. Additionally the…
I don't believe AGI is possible but if it was and it was as subjective as you say what is and isn't conscious, then it starts to take on an even more altogether evil character.
Akin to cloning slave humans or something for free cheap labor.