Earlier quoted context omitted.
Doing basic copyright analyses on model outputs is all that is needed. Check if the output contains copyright, block it if it does. Transformers aren't zettabyte sized archives with a smart searching algo, running around the web stuffing everything they can into their datacenter sized storage. They are typically a few dozen GB in size, if that. They don't copy data, they move vectors in a high dimensional space based…
So then copyrighted content scraped is not needed for training? Guess I missed AGI suddenly appearing that reasoned things out all by itself.
GPT-5 is behind schedule
781–790 of 1001 posts
Re: GPT-5 is behind schedule
#782So the team I lead does a lot of research around all the “plumbing” around LLMs. Both technical and from a product-market perspectives. What I’ve learned is that for the most part that AI revolution is not going to be because of PHD-level LLMs. It will be because people are better equipped to use the high-schooler level LLMs to do their work more efficiently. We have some knowledge graph experiments where LLMs contin…
> a revolution in knowledge-worker productivity. That's a nice euphemism for "imminent mass layoffs and a race to the bottom"...
Re: GPT-5 is behind schedule
#783I found that amusing and figured then that we might already be beyond peak AI!
If AI is trained on sources that include The Sun, it's going to end up being garbage and/or constantly needing to be fact checked.
Re: GPT-5 is behind schedule
#784Earlier quoted context omitted.
When you think about it it's astounding how much energy this technology consumes versus a human brain which runs at ~20W [1]. [1] https://hypertextbook.com/facts/2001/JacquelineLing.shtml
It’s almost as if human intelligence doesn’t involve performing repeated matrix multiplications over a mathematically transformed copy of the internet. ;-)
Re: GPT-5 is behind schedule
#785Earlier quoted context omitted.
the context limits on google are nuts! Being able to pump 2 million tokens in and having it cost $0 is pretty crazy rn. Cline makes it seamless to switch between APIs and isnt trying to shoehorn their SAAS AI into a custom vscode (looking at you cursor)
>the context limits on google are nuts! Being able to pump 2 million tokens in and having it cost $0 is pretty crazy rn. What's the catch though? I was looking at Gemini recently and it seemed too good to be true.
> When you use Unpaid Services, including, for example, Google AI Studio and the unpaid quota on Gemini API, Google uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products and services and machine learning technologies, including Google's enterprise features, products, and services, consistent with our Privacy Policy.
Re: GPT-5 is behind schedule
#786Earlier quoted context omitted.
Not worth 1B ? Come on man. I see them improving the tool enough for most people willing to pay 50$ a month for a subscription. And for most companies to be willing to pay 300$ per employee. It's perhaps not there yet but I'm sure they'll reach this amount of value for their offering. It remains to be seen what competition will do to the prices though.
The market of people willing to pay $50 a month for OAI vs $0/month for one of the open source LLAMA variants is not large enough to justify their current valuation, imo
Re: GPT-5 is behind schedule
#787So the team I lead does a lot of research around all the “plumbing” around LLMs. Both technical and from a product-market perspectives. What I’ve learned is that for the most part that AI revolution is not going to be because of PHD-level LLMs. It will be because people are better equipped to use the high-schooler level LLMs to do their work more efficiently. We have some knowledge graph experiments where LLMs contin…
> a revolution in knowledge-worker productivity. That's a nice euphemism for "imminent mass layoffs and a race to the bottom"...
Re: GPT-5 is behind schedule
#788Re: GPT-5 is behind schedule
#789Re: GPT-5 is behind schedule
#79025% of the top 1000 websites are blocking OpenAI from crawling: https://originality.ai/ai-bot-blocking I am betting hundreds of thousands, rising to millions more little sites, will start blocking/gating this year. AI companies might license from big sources (you can see the blocking percentage went down), but they will be missing the long tail, where a lot of great novel training data lives. And then the big sites w…
As it stands, OpenAI has a market cap large enough to buy a major international media conglomerate or two. They'll get data no matter how blocked they get.