Anybody know any good APIs or projects that help you get the article content for a given url? There is of course Mercury API, but are there any alternatives?
Actually uses the original Insapaper site-specific extraction rules (they were hidden after Marco sold it I think) which can now be found here: https://github.com/fivefilters/ftr-site-config/ We get contributions from users of Wallabag (good open source alternative to Instapaper/Pocket).
It also uses the original Arc90 Readability code ported to PHP to figure out where the article content is without any knowledge of the site.
We sell the latest versions (AGPL licensed) but older versions go up in our public repository: https://bitbucket.org/fivefilters/full-text-rss