Earlier quoted context omitted.
I'm doing the same right now. Massive scraping plus "gisting" or document summarization. You're pretty much on the right track; half of those papers are industry standards (my browser marked them as "visited" automatically :-)
So which of those would be a good starting point for someone who has no idea about IR?
i'm not worried about IR because I use a good search-engine library: Montezuma, and it's in Common Lisp. It's Java clone is called Lucene :-)