Live data from Hacker News

Guy running a Google rival from his laundry room

fastcompany.com

151–156 of 156 posts

Re: Guy running a Google rival from his laundry room

#151
post #56

Earlier quoted context omitted.

Yeah people buy residential IPs on the black market. They are essentially infected home PCs and botnets.

you can get paid about $0.10/GB in cryptocurrency (at a few GB per month) to run one on your PC. Apparently they also just buy actual connections sometimes. It's not even unethical - it's just two groups of equally bad businesspeople trying to spend money to block the other one.

I've heard a few horror stories... Since the people using residential proxies aren't necessarily always good people

Re: Guy running a Google rival from his laundry room

#153

Nothing new as it has been done before, the concept is simple enough: step 1: indexer, solr/lucene Step 2: crawler of which there are several foss, build one yourself? or you just run yacy which is a combo of the above, hook combine with an oldschool searx instance and you will be granted the title as seeker by the spirit of Fravia+ who was elder of the searchlores!!! Not only will you filter crap made by machine lea…

xapian is easier and faster. No Java memory eater.

I've once built a good company wide search engine with custom crawlers, and result hooks, eg to crazy SAP or other ticket systems. Gmane was also legendary.

Re: Guy running a Google rival from his laundry room

#154

Earlier quoted context omitted.

It's not limited to physical effort. Wikipedia's example has embarassment in place of effort; presumably, money could also work.

I interpreted to mean that using a search engine is “useless or unenjoyable, or experiencing unpleasant consequences...”, with attention given to the last two feelings. And I can't figure out what that has to do with people who like Kagi and why it’s wrong or irritating for them to do so. Granted I’ve been annoyed by similar occurrences with other services, but not to the point of suspecting collusion between the ser…

Switching to a nondefault technology takes effort and switching to Kagi in particular also costs money, which is also effort for the purpose of the psychological effect known as effort justification. Therefore, people would be likely to rate switching to Kagi as a good thing even if it was exactly the same as Google (says the effect). Therefore, people who say Kagi is good find it exactly the same a Google (implies the commenter).

Re: Guy running a Google rival from his laundry room

#155
post #132

Earlier quoted context omitted.

Avoiding GIGO (Garbage In, Garbage Out). This is why we have computer-variants of Library Science and Archeology, Forensic Science and a bunch of other advanced knowledge (not AI, mind you).

I don't see how this applies as its aggregating a bunch of stuff from random crawlers - if you want to crawl a list of actual domains that's generally considered the list of things that could resolve, so seems like a good starting place.

Smashing stuff together by pure probablistic word association like AI do today is really, tsk tsk.

That's why it is important to clearly define words like love, oh wait.

Re: Guy running a Google rival from his laundry room

#156
post #93

Well, I created my own domain index. I have not crawled every page inside domains, but it is not my goal. I have 1542766 domains. Might not be much, but it is an honest work. It is available as a github repo, so anybody that wants to start crawling has some initial data to kick off. Links https://github.com/rumca-js/Internet-Places-Database

Cant you just request the ICANN’s zone files and have the canonical list of the day?

you can, though you must provide a reason compelling enough to the person maintaining access (I provided a few sentences and was approved for most but maybe 20% of registrars declined my request):

https://czds.icann.org/home

also be prepared for thousands of emails about status changes to your access.

Post reply on HN