Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

171–180 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#171
post #97

I think DuckDuckGo is closer to what you want. Same results for everyone, better privacy, and they're proactive about improving their results. https://duckduckgo.com/ Part of the problem is that there's a lot more low-quality content to wade through now than there was in 2005. I think the Google of 2005 would have trouble delivering quality results today also.

both ddg and brave are bing (microsoft) in disguise.

This is not correct. Brave Search owns its own (growing) index and relies on third-parties like Bing for some fraction of the requests. Which is not the same thing as relying fully on Bing or third-parties for results like so many meta-search engines. More detailed answer here: https://search.brave.com/help/independence

Edit: Forgot to say that I work on Brave Search.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#172
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

What we need is a net neutrality doctrine on the server side. Bandwidth is hardly scarce outside of AWS's business model. Ban the crawler user-agent dominance by the big search engine players. "Good behaviour" should be enforced via rate limiting that equally applies to all crawlers, without exemption for certain big players.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#173

Earlier quoted context omitted.

Stemming is not meaningfully a natural language processing technique, any more than arithmetic is a technique of linear equations.

Is it not the processing of natural language?

Would you call addition a system of linear equations?

No, you don't use the college senior label for the highschool freshman topic. You use the smallest label that fits.

It's string processing.

NLP is actually understanding the language. Stemming is simple string matching.

Playing the technicality game to stretch fields to encompass everything you think even marginally related isn't being thorough or inclusive; it's being bloated, and losing track of the meaning of the term.

Splitting on spaces also isn't NLP.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#174

Natural Language Processing is a pox on modern search engines. I suspect that Google et. al. wanted to transform their product into an answer engine that powers voice assistants like Siri and just assumed everyone would naturally like the new way better. I can't stand how Google is always trying to guess what I want, rather than simply returning non-personalized results solely based on exactly what I typed in the tex…

I doubt they assumed it was better. I expect they did a ton of user testing and found that it was better for most people. And I'm sure it is. HN users are very much a niche audience these days.

Right. Bing switched to this method as well, as did Facebook, Twitter, Amazon, and pretty much every other company that has the ML resources to do this. They obviously had a good reason to do so, beyond assumptions.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#175

I'd like to see a "just search" engine, all it does it search for a specific string, case insensitively, across the entire web. No curation or anything, just sorted in lexicological order closest match first maybe falling back to page age if it has more then one exact match. Perhaps give me some regular expressions as well.

Maybe a “stability factor” could be calculated. Whereas earlier new content was king, I now value a stable long term source of information. So domain age + page age + content variability + dependency on ads. That might give more honest sources a go.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#176
I do like the idea that instead of crawling and indexing, the next generation search will likely be more like a federated community search app that indexes the stuff members actually read. Google search isn't so much a repository as a consensus about what's important, hence why it's so politicized to the point of becoming unreliable, but also why it too is vulnerable to disruption.

Imo, 2005 google got initial traction because of its tech forum post indexing, as I remember my switch to it was because it became an extension and then replacement for manpages. In that sense, what made it good was it reflected the consensus of what its incredibly influential userbase thought was important and just managed that really well. The demographic impact of the U.S. Gen X all using it at once didn't hurt either.

The equivalent today, as a lot of us say, is that blockchains are in the 1997 internet phase, and the service that makes the content of those as navigable as the 90's internet, will likely grow in a similar way.

Search that provides young people with privacy and freedom to pursue their true interests will be the dominant strategy. Its success will be because it's a product that rides growth, and not because it "solved a problem." Imo, we all index too much on the privacy pattern because the freedom pattern is too risky.

What's changed since that time are the maturity of things like Bloom and other probabilistic filters, Apple's private set intersection, differential privacy, zksnarks, and everybody you'd ask an opinion from now gets their content through mobile devices. Apple's ecosystem is equipped to do this kind of search, but they're too exposed politically to get into it. Meta will likely go there, but nobody's going to trust them willingly.

A protocol that generated a cryptograpically strong anonymous index from your browsing - and instead of putting it on google's servers, it was on a chain, or the content index information and its evolving consensus score was included in something like a DNS record - may still unseat these ensconced interests. IPFS and other P2P or torrents might do something like that as well. Blockchains maybe good for that consensus/desire score.

It's not something you architect and design top down that has to solve all cases, it will be just another useful product that grows while riding a demographic change. It would be on the level of inventing HTML/HTTP again, which, when you think about it, was just another dude making a thing he needed.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#177

Earlier quoted context omitted.

Let's imagine I want to talk to the author of the content. How can I do that if it's just a souped up markov chain?

The markov chain can also power a chat bot.

But then they would need to know that the person sending the email is the same person that read a specific article.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#178
post #42

Also, where are the books about writing a search engine? Knuth's "Searching and Sorting" volume desperately needs an update.

I don't even know if anybody has written a book specifically about search at "web scale" (no MongoDB jokes here, please). But about the closest things I know of would be something like:

https://www.amazon.com/Managing-Gigabytes-Compressing-Multim...

https://www.amazon.com/Information-Retrieval-Implementing-Ev...

https://www.amazon.com/Introduction-Information-Retrieval-Ch...

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#179

I'd like to see a "just search" engine, all it does it search for a specific string, case insensitively, across the entire web. No curation or anything, just sorted in lexicological order closest match first maybe falling back to page age if it has more then one exact match. Perhaps give me some regular expressions as well.

That would be easily the worst search engine ever deployed. Imagine just returning all docs containing the word “bicycle” in chronological order. Useless.

For "Bicycle" it would suck but I don't often use search engines that way, for "High Timber ALX 29" you'd probably get something like this: https://www.schwinnbikes.com/products/high-timber-alx-29?var...

I wouldn't use it for everything but sometimes that is the exact behavior that I want. I'd use duck duck go for more general searches.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#180

Everyone runs in the other direction anytime a search engine is mentioned. The thought of competing with Google turns people off. Even in 2021, despite how bad it's become, it's still miles ahead of other competitors.

I disagree. A lot of people I know already switched to Duckduckgo. Google’s ability to get relevant results is dropping like a brick, while the quality of DDG has been improving slowly but steadily.
Post reply on HN