Hopefully some day soon the internet will be searchable again.
Thanks to everyone involved in attempting to make this happen (preferably in a non-profit-maximized way).
(Said before at https://news.ycombinator.com/item?id=32034390)
121–130 of 396 posts
Hopefully some day soon the internet will be searchable again.
Thanks to everyone involved in attempting to make this happen (preferably in a non-profit-maximized way).
(Said before at https://news.ycombinator.com/item?id=32034390)
The biggest challenge with making a search engine is to combat adversarial SEO. It's an issue that's very easy to be overlooked when you are small, but at Google scale, your enemies have billions of dollars to make from your visitors. I bet Google spends at least as much to combat that, and it's extremely hard to deal with while being open-source. It's useless to call for a non-profit search engine without tackling t…
Perhaps a better approach would be building an open source www index or even a full current cache - as an enabler for people to build their own search engines? Right now it is extremely difficult to build your own web crawler that would compete with Google. And that is not because of the technology, but because multiple sites will prevent your bot from accessing them if you're not Google or Bing - either through robo…
The biggest challenge with making a search engine is to combat adversarial SEO. It's an issue that's very easy to be overlooked when you are small, but at Google scale, your enemies have billions of dollars to make from your visitors. I bet Google spends at least as much to combat that, and it's extremely hard to deal with while being open-source. It's useless to call for a non-profit search engine without tackling t…
Are your sure they're combating it? It seems like they've given up.
It's unfortunate that with all the immense value that search engines provide the idea of paying a small monthly or annual fee to use a search engine is incomprehensible for most people.
Quite a neat way to crawl websites using a browser extension. That by itself is a form of donation to the search engine. Maybe in the future you can have dedicated software for self-hosted clients that users can run to crawl and index websites for mwmbl? Kinda like folding@home. How are the batches of URLs to be crawled generated/discovered and posted at your API? How do you deal with duplicate crawls?
It changed a bit in the implementation.
Earlier quoted context omitted.
I think this is a great idea. How does this work with copyright? Search engines seem to be able to download a reproduce content from scraped pages (and wrap it in ads, and derive content from it) this is called “indexing” when they do it but scraping when everyone else does it.
> and wrap it in ads, and derive content from it i am probably missing something but can you give an example where this happens?
(I don't know if this happens for this specific example, but Google does this for some searches)
Hope this takes off. Also hope he works on his math when it comes to funding, 1% of 40 billion is 400 million, not 4 :p