Live data from Hacker News

Internet Search Tips

gwern.net

61–70 of 81 posts

Re: Internet Search Tips

#61

I've seen two links to this domain in as many days and I absolutely can't stand the developer's treatment of links. I can't actually click on them, because just hovering them opens them in some ephemeral pseudo-popup, and I can't read the main article, because scrolling will invariably hover over some link which will then open that popup and cover up the article I was reading. It's absurd to me that this apparently i…

I actually like it, I just wish it wasn't so fast. The delay is less than a second when really it ought to be more like a a second and a half.

Re: Internet Search Tips

#62
post #51

To augment the suggestion around learning / using hot-keys, ctrl-shift-t (chrome) undoes the most recent tab or browser window close. It's insanely handy, but from casual observation, not well known / used. If you have a mouse with some additional side buttons (that you don't already have mapped) I can strongly recommend mapping them to ctrl-pgup and ctrl-pgdn, so you can go left/right with your tabs in the current b…

This very closely mirrors my experience. I bought an MMO mouse to use in Photoshop but found that it was probably even more useful after mapping forward, back, switch/kill/resurrect tab and so on. I also configured normal click to “always open in the same tab/window” and open in new tabs exclusively with middle click. The efficiency increases are impressive.

Gwern’s tips for “quick searching” (which I implement using DDG !bangs and Firefox Saved Searches) are very useful too, when you know you’ll use a website regularly, I have about 30 custom saved searches now.

Re: Internet Search Tips

#63
Gwern mentions the benefit of being affiliated with a university so you can use ILL (in particular, Illiad). I think it's worth mentioning some ways you can get ILL without having to be a student or professor. (Note that these are not actually going to be practical for most people, but for people who happen to be in the right situation to take advantage of them, they can be quite useful.)

1. Live in New York City. Yes, the New York Public Library seems to provide access to Illiad! I haven't actually tried making use of this, but it's on the website. Obviously you're not going to move to New York just to take advantage of this, but if you happen to already live there, you have this option!

I expect there are other cities and non-university organizations that provide access to Illiad; I mention New York just because it's one I know of. If other people know of others, I'd be glad to know of them!

2. Get a position as a "visiting scholar" at a nearby university. :) OK, this one will likely require knowing someone there, and maybe having a PhD since there may be some minimum requirements, but generally if you can get a professor to name you can one you can become a "visiting scholar" -- this isn't a job, they don't pay you anything and you're not required to do anything for them, and as such there's no hiring process, you just can get named one if you meet the minimum requirements (they may want to see some sort of CV also). They won't pay you any money but you will get library access, including ILL! So, y'know, that's useful. :)

Obviously that route isn't open to everyone either. But depending on your situation it can certainly be easier than enrolling as a student or getting a job as a professor!

Re: Internet Search Tips

#64
post #51

To augment the suggestion around learning / using hot-keys, ctrl-shift-t (chrome) undoes the most recent tab or browser window close. It's insanely handy, but from casual observation, not well known / used. If you have a mouse with some additional side buttons (that you don't already have mapped) I can strongly recommend mapping them to ctrl-pgup and ctrl-pgdn, so you can go left/right with your tabs in the current b…

This very closely mirrors my experience. I bought an MMO mouse to use in Photoshop but found that it was probably even more useful after mapping forward, back, switch/kill/resurrect tab and so on. I also configured normal click to “always open in the same tab/window” and open in new tabs exclusively with middle click. The efficiency increases are impressive. Gwern’s tips for “quick searching” (which I implement using…

Back when it still worked for most of the web (and before it became the progenitor of most contemporary browsers) KDE's Konqueror was my favourite and default browser.

It came with a bunch of built-in web search keywords that sped up your search intents (all customisable of course). Usually two or three letter prefixes, that were terminated by a colon, so you could, f.e. 'ggl: khtml history' (takes you direct to first google 'lucky' result) or 'wp: charles eaton' (search and show wikipedia's top hit for charles eaton), etc.

Re: Internet Search Tips

#65
post #45

Some neat stuff here. Only VERY briefly mentioned (so briefly I missed it at first) however: Substituting Yandex for Google is great for many use cases. Being Russian, Yandex is no doubt heavily censored, but only for things important to Russian politics . Ironically, this means that for non-Russian users, it's considerably LESS censored than Google, which has SEVERELY crippled its search in recent years in the name…

Is there a search portal that does backend-side searches of all these politically-disjoint large search providers, and then merges and deduplicates the result?

I have actually been working on a prototype for something similar to the "meta-search" engines of the 90s, but the intent is to deal with the present situation of search engines placing limits on number of results, and to rid SERPs of needless cruft (Javascript, CSS, advertising, etc.) making them look more standardised according to personal taste. I never really liked meta-search engines because they were too slow to return results. Instead this script lets me query search engines directly, one or more at a time, in succession, or all at the same time, and merge the results into simple, aesthetically-pleasing HTML files.

I use a text-only browser to read HTML so what I am describing here is not designed with "modern" browsers in mind.

The approach I take to avoid search result limits is that I search from the command line and store the result URLs in standardised "search results files" in a "search directory", one file per unique query. When I reach the results limit for the particular search engine, I can repeat the search on another search engine. The search results file is created for two reasons: 1. it allows me to strip out all the cruft from SERPs and mix results from different search engines into one file that looks great in a text-only browser, and 2. it tells me where I left off for each search engine, so I can continue any search at a later time, i.e., get more results.

The script reads from the search directory and presents me with a menu of numbered searches. I continue a search at any time by selecting a number. For example:

   1 this is an example  
   2 foo
   3 foo bar
   4 foo bar baz
I typically browse results by pointing the text-only browser at the search directory.

The search result files are each named according to the URL-encoded search string. The file format is very simple. There are three type of lines: 1. a title tag for the search query, 2. a link for each result URL, and 3. an HTML comment for each HTTP request, indicating to the script where I left off, i.e., the last result number; this is more or less equivalent to a "Next page" or "More results" link. Each result URL and comment is prefixed with a search engine identifier, i.e., a prefix. Thus the script can read the search results file comments and I can easily see which results came from XYZ search engine versus ABC search engine.

  this is an example
  
  X https://example.com
A https://example.net/index.html
This example search results file above shows 1 result retrieved from XYZ search engine (prefix "X") and another result retrieved from ABC search engine (prefix "A"). It shows comments indicating where to continue the search for each search engine. Normally there would be around 50-100 result URLs per HTTP request.

This approach assumes the search result limits are temporal, i.e., they are limits on how many results can be retrieved in a given period. That may not be true for every search engine.

Re: Internet Search Tips

#66

Unfortunately, using the advanced search operators "too much" can get you banned from Google for a few hours, where you get an infinite series of CAPTCHAs. What counts as too much seems to vary widely, but I've triggered it with as few as one query for some obscure phrase using site: . Google is definitely far worse for obscure things than it was a few years or a decade ago. 2010 is roughly when I started noticing it…

Yet another reason to employ DDG as your primary search engine. I've never been rate-limited by it.

(Google Web Search rate limits seemed to kick in around 2015 or so.)

You can still search Google (or metasearch) with bang queries, so, !S (for the Startpage Google proxy search) or !G (for google directly).

Google rate-limiting / CAPTCHA is vastly worse if you're on Tor or VPN. For the most part I simply avoid Google entirely.

Re: Internet Search Tips

#67

Some neat stuff here. Only VERY briefly mentioned (so briefly I missed it at first) however: Substituting Yandex for Google is great for many use cases. Being Russian, Yandex is no doubt heavily censored, but only for things important to Russian politics . Ironically, this means that for non-Russian users, it's considerably LESS censored than Google, which has SEVERELY crippled its search in recent years in the name…

I haven't used Yandex much, But I started noting down failed search results according to the search engines in the hopes of one day creating a 'Search Engine Wall of Shame' to provide genuine feedback to the Search Engines in the areas they could improve[1]. I'll try Yandex for such queries too.

[1] Those who are interested in a 'Search Engine Wall of Shame', URL for the discussion in my profile(#207).

Re: Internet Search Tips

#68

I've seen two links to this domain in as many days and I absolutely can't stand the developer's treatment of links. I can't actually click on them, because just hovering them opens them in some ephemeral pseudo-popup, and I can't read the main article, because scrolling will invariably hover over some link which will then open that popup and cover up the article I was reading. It's absurd to me that this apparently i…

you can disable those pop-ups.

make one of them appear and click on the gear icon top right.

Re: Internet Search Tips

#69
An excellent set of (re)search methods, many of which I'm well familiar. A few notes and additions:

- There are numerous public-domain full-text archives, including Project Gutenberg, Internet Archive, and many small specialised library collections (usually focused on a given topic, e.g., Online Library of Liberty). Less useful for post-1925 materials, but often high-quality renderings (either scans or proofread re-typeset / typed-in documents) available. Google Books also allows full PDF downloads for public-domain works, generally.

- NYPL's Secretly Public Domain project has been reviewing copyright renewal records to find works published since 1923, and before 1964, whose copyright was never renenwed. Other projects (Internet Archive notably) have been flagging these works as being in the public domain, and hence freed of any download restrictions.

https://www.nypl.org/blog/2019/05/31/us-copyright-history-19...

https://www.nypl.org/blog/2018/03/30/unlocking-record-americ...

- OpenLibrary / Internet Archive increasingly have current under-copyright books available for at least 1hr and up to 14 day loan. The reader is less elegant than it had been in past, but is viable.

- HathiTrust is all but useless with its download restrictions. It's helpful to determine if records exist.

- Worldcat gets only a brief mention by Gwern. It's a union catalog (a combined library catalog of a vast number of libraries worldwide), of books, articles, and other document types, and is an excellent way of determining if a book exists, what an author's output is, and/or the documents within a given search space. !worldcat DDG bang search, "ti:" is title, "au:" is author, "kw:" is keyword. Space any colons (":") occurring within search terms, or omit them entirely. You'll still have to either find the digital record elsewhere, or track down a library, but quite useful.

- You can save online materials to the Internet Archive using the 'save' URL:

   https://web.archive.org/save/
So to save this particular HN discussion we'd specify:

   https://web.archive.org/save/https://news.ycombinator.com/item?id=26847596
You can submit that through any HTTP client (curl, wget, lynx, w3m, your GUI browser, etc.). Requests can be trivially scripted and batched.

This is ... documented somewhere (I stumbled across it myself), though I'm not finding the specifics. Related "save page now" functionality is mentioned here: https://blog.archive.org/2019/10/23/the-wayback-machines-sav...

- Motorised paper cutters are available at some photocopy shops. Inquire as to whether or not you can have books debinded by them. (Generally anything resembling paper is fine, though the blades can be damaged by metal or other materials.)

Post reply on HN