Live data from Hacker News

Google Declaring War on the Web

tante.cc

201–210 of 466 posts

Re: Google Declaring War on the Web

#201
post #190

Earlier quoted context omitted.

Well, pirated. Piracy and stealing aren't the same thing. Regardless, I acknowledged the general issue. However I pointed out that doing so was not a technical necessity. If you base your worldview or actions around X implying Y but then it turns out that actually Y was merely a matter of convenience you're probably going to arrive at a wrong conclusion. There's also the issue where you're emphatically calling it ste…

It absolutely is a technical necessity. You could build a model from scratch today without doing the same thing. And every model attempting to train on AI generated output degrades into nonsense almost immediately. There’s a reason Reddit is making millions of dollars letting these companies mine their human generated content. You think OpenAI or anyone else would pay for that if they could just cyclically train on A…

> attempting to train on AI generated output

I said nothing about that. Good synthetic data does not (typically) involve ML algorithms. Although that might be changing.

I'll politely suggest that you go read the literature before engaging further.

Reddit, Twitter, and similar are valuable because the data covers current events. Their content makes up a reasonably comprehensive timeline of the world at large. You don't need that to train a barebones functional model but it's certainly useful in order to train a knowledgeable one. Regardless, if they're charging for access it clearly isn't piracy so it doesn't seem like your original objection would hold any water in that case.

Re: Google Declaring War on the Web

#204
post #150

Earlier quoted context omitted.

The problem is that it is increasingly difficult to survive as a small business (due to constantly increasing compliance/regulatory/legal burdens), so it makes sense to ‘sell out’ as soon as possible (or just give up early). The rate of small businesses growing into large ones has been decreasing for at least 20 years.

This is tech we are talking about. There are very little, if any, regulatory burdens in place here. The only things that DO hurt SMBs across the board are things like paying for private health insurance and retirement plans. Two core things every worker needs but only massive corporations can truly provide. It's why things like medicare for all and universal childcare are so popular among workers, also why things lik…

>"There are very little, if any, regulatory burdens in place here."

Speaking as someone who works in a small company that designs and manufactures embedded devices, I can tell you that many of the 'minor' regulations which are not supposed to burden small businesses actually do. My (single) biggest annoyance is the conflict minerals reporting requirements which were supposed to apply to very large companies, but have been 'passed down' to smaller suppliers (as anyone with half a brain would have expected). There are many other KYC, CBP, and other regulations which have substantial impacts as well.

Re: Google Declaring War on the Web

#205

I don't understand the endgame here. Websites let Google crawl their content in exchange of traffic. If Google cuts that out completely, what incentive do websites have to not block the Google crawlers? I understand that Google is feeling an existential threat from other AI products that provide answers directly. But they must also understand their symbiotic relationship with the web.

If they block Google’s crawlers no one visits their site ever.

Re: Google Declaring War on the Web

#207
post #94

I don't understand the endgame here. Websites let Google crawl their content in exchange of traffic. If Google cuts that out completely, what incentive do websites have to not block the Google crawlers? I understand that Google is feeling an existential threat from other AI products that provide answers directly. But they must also understand their symbiotic relationship with the web.

What I really don't understand is where the next generation of training material will come from. If websites stop being published and/or crawled, how will the machine continue to be fed.

Probably real life. At some point, these LLMs are going to be good enough to just train themselves off of cameras and audio recordings of people out in the real world. They’re going to have robots everywhere constantly listening to what people are saying.

Alternatively, they’re probably betting on being able to get the AGI with everything we already currently have and at that point further training doesn’t matter.

Re: Google Declaring War on the Web

#208
post #136

Earlier quoted context omitted.

An open way to trade, store, and export lists of websites in a way that works seamlessly on desktop and mobile browsers would be pretty neat.

Like bookmarks and links?

On a higher level than individual URLs and separate from browser favorites. Something like versioned packages of links with decorations.

Something like a ".urlpackage" format that will have

- a list of urls

- optional metadata for each url, such as image, description, last-known-good

- metadata for the entire package, including version, an image, a favicon, and a description for the entire package that a client could use to present it nicely to the end user.

It'd be cool if my phone could open this format, show me the image and description with the list of links, and let me browse them, add them to my bookmarks, or add to the collection and make a new .urlpackage that I could then share back or publish somewhere.

It's probably possible to simply do this with a self-contained HTML file or similar I guess, though.

Re: Google Declaring War on the Web

#209
post #85

I feel like AI has gotten to the point where the message is: If you want to make something (art/code/music/writing) you can do it for your own enjoyment, but you aren't allowed to make money from it anymore; only the large corporations can make money from content. If you do release something creative, it'll just be fed back into the machine to be copied over and over.

At least for art - I don't think you'll find anyone who actually enjoys art hanging up anything produced by AI on their walls. For these kinds of "customers", they could equally easily frame & hang up a poster of the Mona Lisa. Artists are not at threat, if anything, AI makes original artworks more precious & enjoyable.

I think this is only true in a vague and abstract way. In reality, AI devalues labor (in general) and the worth of artists (in specific).

Good art requires good patronage and institutional support in turn. No one will have time to produce the next Mona Lisa if they're barely able to make end's meet working a slavish factory job. That's doubly true when the vocations that supported artists—either antiquated, modern, or contemporary (painter, typesetter, graphic designer, etc.)—vanish because AI can do "just about as well."

Art isn't just a divine presence gracing the souls of those deemed most worthy, it's a collection of skills and knowledge that must be built by community over decades of struggle.

On top of the generation of slop, AI is removing some of the final protections that hold these pillars up. That is what should keep us up at night.

Post reply on HN