Live data from Hacker News

On the Future of 12ft

12ft.io

41–50 of 113 posts

Re: On the Future of 12ft

#42
post #35

> How does it work? > The idea is pretty simple, news sites want Google to index their content so it shows up in search results. So they don't show a paywall to the Google crawler. We benefit from this because the Google crawler will cache a copy of the site every time it crawls it. > All we do is show you that cached, unpaywalled version of the page. This seems like something that could be accomplished entirely clie…

You can't do it client side. Well-implemented paywalls clip the content server-side. If you're Google, you get the full HTML, if you're not, you don't.

> If you're Google, you get the full HTML, if you're not, you don't. reply

This company isn't Google either, though.

Or are you suggesting that they're using Google Compute Engine and therefore get the unpaywalled version?

Re: On the Future of 12ft

#43
post #22

Earlier quoted context omitted.

Then search engine crawlers should get paywalled too. The motivation is correct. You run a search on google and get mostly paywalled content. I'm fine with news sites requiring subscriptions to view their articles but they shouldn't also get the benefit of being listed at the top of search results for key terms. Alternatively, the search list should show if content is paywalled or give you search options to remove pa…

I'm not following your logic here. The Economist wants to charge readers because they produce high quality content that's worth paying for (in their estimation). This seems entirely orthogonal to whether they blacklist a web crawler - a crawler they didn't even ask for, which would be all over their website whether they want it or not. I think you're confused because the crawler and the browser both use the same chan…

> I think you're confused because the crawler and the browser both use the same channel and the same protocols to access the information (the website over HTTP). But that's just a detail. Google could send them a hand written form for them to fill out with details of each of their articles and some thumbnail images to be manually entered into a Google database for all we care.

I disagree; I care quite a bit about whether the Google results are about the actual page I'm going to see or about what the page author claimed the page would be about. (Indeed I'm old enough to remember that what originally set Google apart from competing search engines was that it would ignore the meta keyword tags that authors used to describe their pages, in favour of indexing the visible page content directly)

Re: On the Future of 12ft

#44

From the home page explanation.... > How does it work? > The idea is pretty simple, news sites want Google to index their content so it shows up in search results. So they don't show a paywall to the Google crawler. We benefit from this because the Google crawler will cache a copy of the site every time it crawls it. > All we do is show you that cached, unpaywalled version of the page. — https://12ft.io/ They're just…

I'm also confused. Their description seems incomplete, perhaps intentionally so?

I wonder if they're accessing pages from Google Compute Engine in an attempt to appear as a legitimate Google crawler?

If they're just loading the cached Google pages like anyone else, I don't understand why it's hitting their servers at all.

Re: On the Future of 12ft

#45
post #36
post #35

Earlier quoted context omitted.

You can't do it client side. Well-implemented paywalls clip the content server-side. If you're Google, you get the full HTML, if you're not, you don't.

But I could change my clients User Agent to match the Google Crawler?

Sure, and then you've got a cat-and-mouse game with all the other ways that sites can fingerprint you.

Re: On the Future of 12ft

#46
post #23

Earlier quoted context omitted.

Someone better tell every movie studio on Earth that's it's disingenuous to make movie trailers.

Surely you don’t need someone to point out to you that those aren’t the same things.

You're right. One of them is professionally produced content, advertised with a preview to whet your appetite, where you're expected to pay for the right to consume it fully. While the other is.... uhh.... the same thing.

Re: On the Future of 12ft

#47
post #24

Earlier quoted context omitted.

I'd be fine with that if they didn't let search engines crawl their content. If they want to be on the web, they have to play by the rules of the web. Instead what they want is to reap the benefits of the web while refusing to participate in the open web.

What the heck are the “rules of the web”? Only valueless information is allowed to be indexed? Are you offended that when you search for a book on Amazon you have to pay to read the whole thing?

[deleted]

Re: On the Future of 12ft

#48
post #22

Earlier quoted context omitted.

Then search engine crawlers should get paywalled too. The motivation is correct. You run a search on google and get mostly paywalled content. I'm fine with news sites requiring subscriptions to view their articles but they shouldn't also get the benefit of being listed at the top of search results for key terms. Alternatively, the search list should show if content is paywalled or give you search options to remove pa…

I'm not following your logic here. The Economist wants to charge readers because they produce high quality content that's worth paying for (in their estimation). This seems entirely orthogonal to whether they blacklist a web crawler - a crawler they didn't even ask for, which would be all over their website whether they want it or not. I think you're confused because the crawler and the browser both use the same chan…

I’m not confused and I don’t really understand why you think I am.

I believe the content that is indexed is the content you can see. Sites used to be penalised, heavily, for returning different content to google. Hiding the paywall for google falls into that bucket.

At a minimum the search results should display if they’re paywalled and provide tools to exclude that content from results.

Re: On the Future of 12ft

#49
post #38

I don't understand. Isn't this stealing?

No, because when you read an article a copy is made and sent to you. You are not depriving anybody from anything.

It is theft under federal law, punishable by up to 5 years in prison.

https://en.wikipedia.org/wiki/No_Electronic_Theft_Act

Re: On the Future of 12ft

#50
post #8

The “Why?” section of the home page is rather disingenuous, since the Economist is not “SEO optimized garbage”, nor does it want you to “sign up for some newsletter”. It’s just an excellent newspaper that needs to pay the staff that writes all those articles. Just like 12ft.io needs to pay its hosting provider.

I can see Google offering “private indexing” to big publications for a fee in the future. Keeping 12ft ladder out of the loop effectively.
Post reply on HN