Live data from Hacker News

Sacrificing accessibility for not getting web scraped

tilschuenemann.de

31–40 of 44 posts

Re: Sacrificing accessibility for not getting web scraped

#31
post #21

That's cool. Hopefully you never post any remotely interesting, because in my very human 2010's way of doing things, I cannot even select and copy some text to my personal notes. This goes well beyond accessibility and bots. I guess the Reader mode, a basic web browser feature meant precisely to read articles, wasn't an expected use case either?

Simple solution: Screenshot it, then ask your AI of choice to read that:) (Or use any other OCR solution you like; I've got a prototype that takes a screenshot and runs it through tesseract.)

That's funny :)

We'll come full-circle when web authors provide custom-made glasses to decipher their sites, as the plain rendering will be obfuscated to prevent OCR too.

Re: Sacrificing accessibility for not getting web scraped

#32
post #3

Congratulations, I guess? I can't read your content. But ... The machines can't either, so ... great job! Although... Hmm! I just pasted it into Claude and got: When text content gets scraped from the web, and used for ever-increasing training data to improve. Copyright laws get broken, content gets addressively scraped, and even though you might have deleted your original work, it might must show up because it got c…

Does (I’m assuming) your screen reader cope with text that’s displayed in (for example) a raster or a vector image?

Re: Sacrificing accessibility for not getting web scraped

#33
post #18

At least the blog author is self-aware about making accessibility worse? I just found it funny how reactionary and backfire-y this was. (In politics, a reactionary is a person who favors a return to a previous state of society which they believe possessed positive characteristics absent from contemporary society.)

I sympathize with the reactionary. Obviously there's no putting the genie back in the bottle, but it would be nice to live in that world where writing stuff helped human readers more than it helped billion-dollar corporations.

Re: Sacrificing accessibility for not getting web scraped

#34

I have a dumb question: What if we put a simple password in front of every website that everyone knew, like "password". Upon click of login, the user agrees to the terms of service which exclude all automatic scraping. I know this is a dumb idea, but I would love to know exactly why.

Law is only as useful as your ability to enforce it. Would you have the legal means (and will) to go after a random foreigner who still scraped ypur site after accepting the terms? Multiply by 1,000 times a day.

Re: Sacrificing accessibility for not getting web scraped

#35

I have a dumb question: What if we put a simple password in front of every website that everyone knew, like "password". Upon click of login, the user agrees to the terms of service which exclude all automatic scraping. I know this is a dumb idea, but I would love to know exactly why.

Legal agreements won't stop companies who don't care about legal agreements, they'll just add "popup.write('password').submit()" and move on. Multiply by 1000x, and you've solved nothing, while making the UX for normal users worse.

I should have been more clear with the goal.

I know it's technically very easy to get around, but would it give the content owner any stronger legal footing?

Their content is no longer on "the open Internet," which is the AI labs' main argument, is it not?

Re: Sacrificing accessibility for not getting web scraped

#36
I mean, the core problem is the same as DRM: you want to distribute your data without actually distributing it. The endgame is simply rendering and OCRing, which will likely be feasible at scale, but in the meantime, you've cut off a bunch of your legit audience, broken RSS, broken search engine indexing, and AI can still scrape it when requested.

Re: Sacrificing accessibility for not getting web scraped

#37
post #14

Jeez, all this "LLM avoidance" is so horribly silly. We will remember it in 10 years like we now remember the decision to setup our great European cookie popup laws. As well intentioned, but very detrimental. You will not stop scrapers. Period. They will just pay for a service like firecrawl that will fix it for them. Here in Poland one of the most notorious sites implementing anti-bot tech is our domestic eBay compe…

> You will not stop scrapers.

That's no reason to go down without a fight!

> Second argument against these "protections" is, there are people behind bots

That doesn't hold much water in the context of hosting a blog.

Re: Sacrificing accessibility for not getting web scraped

#39
post #3

Congratulations, I guess? I can't read your content. But ... The machines can't either, so ... great job! Although... Hmm! I just pasted it into Claude and got: When text content gets scraped from the web, and used for ever-increasing training data to improve. Copyright laws get broken, content gets addressively scraped, and even though you might have deleted your original work, it might must show up because it got c…

Android's built-in OCR worked for me. I was able to copy the text.

Re: Sacrificing accessibility for not getting web scraped

#40
post #31

Earlier quoted context omitted.

Simple solution: Screenshot it, then ask your AI of choice to read that:) (Or use any other OCR solution you like; I've got a prototype that takes a screenshot and runs it through tesseract.)

That's funny :) We'll come full-circle when web authors provide custom-made glasses to decipher their sites, as the plain rendering will be obfuscated to prevent OCR too.

I can't wait for the entire website to be a Magic Eye image.
Post reply on HN