On the Future of 12ft
41–50 of 113 posts
Re: On the Future of 12ft
#42> How does it work? > The idea is pretty simple, news sites want Google to index their content so it shows up in search results. So they don't show a paywall to the Google crawler. We benefit from this because the Google crawler will cache a copy of the site every time it crawls it. > All we do is show you that cached, unpaywalled version of the page. This seems like something that could be accomplished entirely clie…
You can't do it client side. Well-implemented paywalls clip the content server-side. If you're Google, you get the full HTML, if you're not, you don't.
This company isn't Google either, though.
Or are you suggesting that they're using Google Compute Engine and therefore get the unpaywalled version?
Re: On the Future of 12ft
#43Earlier quoted context omitted.
Then search engine crawlers should get paywalled too. The motivation is correct. You run a search on google and get mostly paywalled content. I'm fine with news sites requiring subscriptions to view their articles but they shouldn't also get the benefit of being listed at the top of search results for key terms. Alternatively, the search list should show if content is paywalled or give you search options to remove pa…
I'm not following your logic here. The Economist wants to charge readers because they produce high quality content that's worth paying for (in their estimation). This seems entirely orthogonal to whether they blacklist a web crawler - a crawler they didn't even ask for, which would be all over their website whether they want it or not. I think you're confused because the crawler and the browser both use the same chan…
I disagree; I care quite a bit about whether the Google results are about the actual page I'm going to see or about what the page author claimed the page would be about. (Indeed I'm old enough to remember that what originally set Google apart from competing search engines was that it would ignore the meta keyword tags that authors used to describe their pages, in favour of indexing the visible page content directly)
Re: On the Future of 12ft
#44From the home page explanation.... > How does it work? > The idea is pretty simple, news sites want Google to index their content so it shows up in search results. So they don't show a paywall to the Google crawler. We benefit from this because the Google crawler will cache a copy of the site every time it crawls it. > All we do is show you that cached, unpaywalled version of the page. — https://12ft.io/ They're just…
I wonder if they're accessing pages from Google Compute Engine in an attempt to appear as a legitimate Google crawler?
If they're just loading the cached Google pages like anyone else, I don't understand why it's hitting their servers at all.
Re: On the Future of 12ft
#45Earlier quoted context omitted.
You can't do it client side. Well-implemented paywalls clip the content server-side. If you're Google, you get the full HTML, if you're not, you don't.
But I could change my clients User Agent to match the Google Crawler?
Re: On the Future of 12ft
#46Earlier quoted context omitted.
Someone better tell every movie studio on Earth that's it's disingenuous to make movie trailers.
Surely you don’t need someone to point out to you that those aren’t the same things.
Re: On the Future of 12ft
#47Earlier quoted context omitted.
I'd be fine with that if they didn't let search engines crawl their content. If they want to be on the web, they have to play by the rules of the web. Instead what they want is to reap the benefits of the web while refusing to participate in the open web.
What the heck are the “rules of the web”? Only valueless information is allowed to be indexed? Are you offended that when you search for a book on Amazon you have to pay to read the whole thing?
Re: On the Future of 12ft
#48Earlier quoted context omitted.
Then search engine crawlers should get paywalled too. The motivation is correct. You run a search on google and get mostly paywalled content. I'm fine with news sites requiring subscriptions to view their articles but they shouldn't also get the benefit of being listed at the top of search results for key terms. Alternatively, the search list should show if content is paywalled or give you search options to remove pa…
I'm not following your logic here. The Economist wants to charge readers because they produce high quality content that's worth paying for (in their estimation). This seems entirely orthogonal to whether they blacklist a web crawler - a crawler they didn't even ask for, which would be all over their website whether they want it or not. I think you're confused because the crawler and the browser both use the same chan…
I believe the content that is indexed is the content you can see. Sites used to be penalised, heavily, for returning different content to google. Hiding the paywall for google falls into that bucket.
At a minimum the search results should display if they’re paywalled and provide tools to exclude that content from results.
Re: On the Future of 12ft
#49I don't understand. Isn't this stealing?
No, because when you read an article a copy is made and sent to you. You are not depriving anybody from anything.
Re: On the Future of 12ft
#50The “Why?” section of the home page is rather disingenuous, since the Economist is not “SEO optimized garbage”, nor does it want you to “sign up for some newsletter”. It’s just an excellent newspaper that needs to pay the staff that writes all those articles. Just like 12ft.io needs to pay its hosting provider.