Live data from Hacker News

Google vs. the Open Web

interpeer.io

21–30 of 207 posts

Re: Google vs. the Open Web

#21
post #9

How does WEI work with non-browsers, like curl or python requests? I was wondering if there is some motive here to monopolize web scraping (especially with respect to harvesting AI training data)?

I mean that’s part of the point. It’s there to exactly lock out scrapers. Or crawlers, for that matter. What a happy little coincidence.

This is just silly, there exist frameworks like selenium that allow you to run any browser of choice and emulate actual user behavior(clicks, keystrokes). If they go further the emulation layer will have to be moved higher, above the virtual machine running the browser for example. The truth is, this has nothing to do with scraping, scrapers will find a way. This is to stop the majority of people from using ad block.

Re: Google vs. the Open Web

#22
post #21

Earlier quoted context omitted.

I mean that’s part of the point. It’s there to exactly lock out scrapers. Or crawlers, for that matter. What a happy little coincidence.

This is just silly, there exist frameworks like selenium that allow you to run any browser of choice and emulate actual user behavior(clicks, keystrokes). If they go further the emulation layer will have to be moved higher, above the virtual machine running the browser for example. The truth is, this has nothing to do with scraping, scrapers will find a way. This is to stop the majority of people from using ad block.

If I understand this proposal correctly, this is exactly to prevent such things. Yes, of course, it’s to prevent people from using ad block. But a nice side-effect is to block crawlers, or frameworks like selenium as well, so they can „serve ads only to real people“. Of course, people will always find a way to crawl. We already have bot farms that are just remote controlled smartphones lined up somewhere. But it makes it harder for everyone who isn’t Google to compete with Google.

Re: Google vs. the Open Web

#23
I'm of the opinion that Google is no longer "Organizing the world's information" but "Stealing the world's information". Their last keynote proved that with all of the AI products. Regardless of this proposal (which is awful) they are a data vampire using all your info for themselves without permission or compensation.

Re: Google vs. the Open Web

#24
post #21

Earlier quoted context omitted.

I mean that’s part of the point. It’s there to exactly lock out scrapers. Or crawlers, for that matter. What a happy little coincidence.

This is just silly, there exist frameworks like selenium that allow you to run any browser of choice and emulate actual user behavior(clicks, keystrokes). If they go further the emulation layer will have to be moved higher, above the virtual machine running the browser for example. The truth is, this has nothing to do with scraping, scrapers will find a way. This is to stop the majority of people from using ad block.

IIRC those require addons to the browser and would fail attestation and get blocked.

Marionette is built into Firefox so that might work, except it would require Firefox to implement this as well so it can prove itself.

Re: Google vs. the Open Web

#25

So, how can this be bypassed in theory? Any ideas? Brainstorming is allowed. EDIT: Saw a few mention two solutions to disable the automatic verification on iOS & macOS. https://blog.cloudflare.com/how-to-enable-private-access-tok... https://support.apple.com/en-us/HT213449

Build a new Web w/o Google.

We have been doing it! Join us!

It isn’t perfect but we are ahead of most others (Mastodon, Matrix). We have spent TWELVE YEARS building the free, permissionless open source platform for anyone to assemble and host their own community software with all the features of Facebook/Twitter/TikTok for their own community:

https://github.com/Qbix/Platform

We are about to roll out version 2.0 — I have never done this before but I would like to invite whoever wants to learn about it or build on it, to a Zoom webinar where I will demo anything and answer any questions. Starting in Q3 this year all the webinars will take place on our own platform — no Calendly, no Zoom, no Google, just the free open Web.

Anyway, sign up here if you want. Will do it every Sunday throughout August:

https://calendly.com/qbix/qbix-2-0-platform-demo

Whether you’re a developer, a businessperson, or just want to learn about the latest technologies moving the Free Open Source Web forward, this platform can help empower you to build and engage a community around yourself and your projects.

Re: Google vs. the Open Web

#27
post #5

The bright side is, when Google really pushes through with this WEI nonsense, it will not only break the Web, it will also create some kind of premium Web run by Google, analogous to gated communities. And then the worldwide Internet ad market bubble will finally burst.

GoogAOL Web 3.0

...send it to the gooGAOL

Re: Google vs. the Open Web

#28

I mentioned this in the other WEI thread and I’ll do it here again: Instead of simply flailing our collective arms around complaining about an evil corporation, has anyone written to the respective competition authorities (such as the FTC in the US or CCI in India) about the potential anticompetitive effects of this proposal?

FTC right now has awful leadership. They only care about blocking mergers and scoring political points.

Re: Google vs. the Open Web

#29
post #21

Earlier quoted context omitted.

I mean that’s part of the point. It’s there to exactly lock out scrapers. Or crawlers, for that matter. What a happy little coincidence.

This is just silly, there exist frameworks like selenium that allow you to run any browser of choice and emulate actual user behavior(clicks, keystrokes). If they go further the emulation layer will have to be moved higher, above the virtual machine running the browser for example. The truth is, this has nothing to do with scraping, scrapers will find a way. This is to stop the majority of people from using ad block.

>If they go further the emulation layer will have to be moved higher, above the virtual machine running the browser for example.

Your hypothetical change of emulation tactics won't work. You're analyzing at the wrong abstraction level.

The "attestation tokens" to validate the integrity of the web browser environment would come from a 3rd-party (e.g. Google Play services).

For example... Today, hacks like youtube-dl work because implementing client-side code to "solve javascript puzzle challenges" is still inside the "world" that Google-server-to-browser-client present to each other. Same for client-side solvers for Cloudflare captchas. The "3rd-party attestation token" breaks those types of hacks.

Re: Google vs. the Open Web

#30
post #28

I mentioned this in the other WEI thread and I’ll do it here again: Instead of simply flailing our collective arms around complaining about an evil corporation, has anyone written to the respective competition authorities (such as the FTC in the US or CCI in India) about the potential anticompetitive effects of this proposal?

FTC right now has awful leadership. They only care about blocking mergers and scoring political points.

Isn’t blocking mergers one of the founding goals of the FTC?
Post reply on HN