Live data from Hacker News

The next Google

dkb.io

441–450 of 561 posts

Re: The next Google

#441
post #324
post #225

Earlier quoted context omitted.

This. People don't realize that the early web was elitist. Now, the entire population is online. And, as you said, most people simply don't care about the stuff we care about. That's also why "Google's search results are soo bad." They're not. For the bulk of Google's visitors, they're good enough.

Indeed. We live in the eternal September. But to be honest the internet is still the internet. The web still exists. Any lamentation at the loss of the "old" internet is that you don't have more angry uncles spewing political rhetoric on your motorcycle forum. Do you want the unwashed masses in your specialist forums? I certainly don't. Seems to me things are working pretty well. I do worry about the next generation…

> Take some youngsters under your wing and show them how a keyboard works instead of a touch screen.

This x1000.

The only way to ensure that future people can continue to maintain things is to teach those who are interested.

While I think we need to scale back the opaqueness of tech to the average user[0], there will always be some who just know more.

[0] By which I don't mean "more configuration options", because the average user doesn't want those.

I mean make it easy to access the source - and just as importantly make it easy to use changes to said source, so those who think, hm, I want this, can learn to fix it themselves, or find someone who can.

Re: The next Google

#442
post #225

Earlier quoted context omitted.

This. People don't realize that the early web was elitist. Now, the entire population is online. And, as you said, most people simply don't care about the stuff we care about. That's also why "Google's search results are soo bad." They're not. For the bulk of Google's visitors, they're good enough.

> For the bulk of Google's visitors, they're good enough I'm not sure that's true. Just to pick an example, basically any type of product search now leads to auto-generated spam/borderline-spam websites that scrape reviews from Amazon, which are themselves often fraudulent. I guess you could say that's "good enough" but only in the sense that someone being scammed at some point consents to something that makes them a…

I've been using Kagi search and I find the same stuff there. I can block them from my personal results as I encounter the garbage, but it's there by default.

And the people trying to get their garbage spam sites to rank high are wise to the common tricks people use to attempt to get better results. Just today I accidentally clicked on a result that from the title looked like it would lead me to reddit. I should have paid attention to the URL because instead I was sent through a bunch of redirects to some terrible spam site.

Re: The next Google

#443

Earlier quoted context omitted.

The death of the web is because most people don't want what you want. They don't mind walled gardens, so long as they are easy to use and have the content and connections that they want to see. The audience of HN is extremely skewed towards preferring systems that allow tinkering but that's not what the market wants.

People don't mind walled gardens because for a while they've been good enough if not better than the previous status-quo. However, those walled gardens are decaying such that there might actually be demand for something better if it existed.

> However, those walled gardens are decaying such that there might actually be demand for something better if it existed.

Some walled gardens, like facebook, are decaying. Others, like tiktok, are vibrant, still fresh and new.

Walled gardens are a natural result of the pursuit of capital. If you burn VC money, but create no moat (it turns out the walls of a walled garden are also a moat, weird huh), then as soon as you introduce a clever algorithm that introduces 50% more ads to extract profit, your users will all leave.

As such, a business is incentivized to build these walls and moats.

How do we avoid this?

Well, we do have examples. Mastodon and other open source projects eschew walled gardens in favor of free software ideals. Web3 embraces a certain "decentralized" vibe which lends itself in this direction. Universities, and other public-ish institutes like DARPA, created the original internet and many of its technologies.

Unfortunately, open source projects will struggle to advertise or find users. They do not have the initial capital to get as much momentum as the VC-funded alternatives. Web3 seems surely doomed to end up also building walled gardens for the crypto-anarcho-vibe is only skin-deep, and many a regular business is now highly involved.

This leaves government entities. The government is the one group that is both well funded enough, and has motive to create protocols which prioritize the user's freedom over profit (after all, the users will pay taxes regardless of how high the walls are).

It seems to be out of fashion these days for the government to actually do anything though, so perhaps there is no more chance of that than of a free software project managing the same.

Re: The next Google

#444
post #430

If memory serves me correctly Google once had a desktop widget with search like neeva. I could search email, files, Web, all at once from task bar. Neat. Maybe I dreamt that.

Yes you are correct. Google Desktop Search. Connected Apps in Neeva is similar but better. It supports indexing more apps, like Figma and GitHub.

Re: The next Google

#445
post #88

Earlier quoted context omitted.

It's actually so good I plan to pay when they start charging.

It's really good in some ways, and very lacking (for me) in others. The search is fine. Good enough that I would switch. However, if I was out and about and quickly needed directions to Walmart, my normal flow with Google is Pull out phone -> Safari -> type "Walmart" -> click on the map -> Maps app opens and starts guiding me When I switched to Kagi, the flow went like Pull out phone -> Safari -> type "Walmart" -> to…

Kagi supports DDG bangs. So just type `Walmart !g` and you will get the google search output

Re: The next Google

#446
post #396

Earlier quoted context omitted.

> Get rid of that noise and your hardware goes a lot longer. What qualifies? What defines signal, what noise? I agree, that a lot (probably nearly all) pages will receive very, very little traffic/search requests. But are these therefore not relevant? > I'm running a search engine on consumer hardware out of my living room that can index 100 million documents. That's extremely cool. I would love to know more. To me a…

You should try looking at people's profiles on HN - just click on the username.

Why? I don't change my reply based on the author. I reply to a statement to the best of my knowledge regardless of the author behind it.

And I learned already a lot in this thread after the explanations unfolded.

The initial statement sounded exactly like the armchair "experts" one so often encounters. Actually this was for a long time the first time that there is a person with substantial experience in the problem space behind such a statement.

Re: The next Google

#447
> For this new generation, privacy is necessary,

Why is this always a given? Yes, privacy is good, but honestly, I find what I'm looking for in the first few links about 99% of the time with Google, because of the lack of privacy. They have 20 years of search history on me, including maps searches. They know where I live and where I go and what I like and what I buy, and they can read all my email.

And I get better search results because of it.

If I search for [haircut], I get the cheap places near me, because Google knows I'm cheap when it comes to haircuts, because they've seen where I've gone before. To get that on an anonymous search engine, I'd have to search for [cheap haircut near $home_address], so now I've had to type extra words and I've given up my privacy anyway, just with extra effort.

I'm not the biggest fan of Google knowing that much about me, but I also know it gets me great results, and that's a tradeoff I'm willing to make (and lots of other people are too).

Re: The next Google

#448

Earlier quoted context omitted.

All you have to do to improve on Google at this point is to do less, make it less bad, i.e. a process of removals, not additions. Just do the same thing, but without all the shitty extra stuff. But then what's the business model? (Since that is in fact most of the shitty stuff. Oh sure there's still the SEO spam, and you're in that arms-race, like it or not, even if you're not a successful search engine, so you do th…

> just do the same thing as google i know what you're saying, but this isn't exactly a triviality

Just scrape google like they scrape everybody else.

Re: The next Google

#450

Earlier quoted context omitted.

> Get rid of that noise and your hardware goes a lot longer. What qualifies? What defines signal, what noise? I agree, that a lot (probably nearly all) pages will receive very, very little traffic/search requests. But are these therefore not relevant? > I'm running a search engine on consumer hardware out of my living room that can index 100 million documents. That's extremely cool. I would love to know more. To me a…

I think I was editing the comment while you were replying. Sorry about that. I was just adding to it though, didn't really rug pull on your response so I think it's fine. > What qualifies? What defines signal, what noise? I agree, that a lot (probably nearly all) pages will receive very, very little traffic/search requests. But are these therefore not relevant? Now this is a proper difficult problem with (probably) f…

> I [...] didn't really rug pull on your response so I think it's fine.

No you didn't. All good. And I learned a lot from the extended answer. So I am thankful for the explanation.

> Developing heuristics for this is a bit of a hobby horse of mine. It feels tantalizingly almost doable with just a little bit more resources and time than I have.

I can totally understand the feeling. There are quite a few things that I'd like to go deeper into either at work or in private. But alas time.

> Now this is a proper difficult problem with (probably) fairly subjective answers.

I agree. And I don't have answers ready. A lot boils down to preference. Personally, for example I prefer written content over video. Except in a few areas were I like (some) explanatory videos. To me it comes down to the question of how easy I can skim the content when I am looking for an answer.

On the other hand - for deep immersion into a topic I use multiple media formats.

In terms of web search I sadly nowadays need to sift through a lot of seo-fied content that is there either to build a (personal) brand or to attract clicks for advertising revenue/affiliate revenue.

So in principle I agree with you on the noise problem. Still I also believe that there are real great gems to be found in the long tail. When I still feel like I came late to the party, but when I started out in the web in '97 there were so many lovely, quirky sites. So many places that people had put a lot of time, energy and thought into. And sites so packed full of information that I came away not only with more knowledge, but in awe that somebody would give this knowledge away for free.

There also were quite a number of horrible sites (my first ones probably included). So there was a noise vs. signal problem back then. Maybe not to the extent today, though.

> The machine it's on is a Ryzen 3900X with 128 Gb RAM. Most of the index is on a single 1 Tb consumer grade SSD.

Call me impressed. Sounds absolutely cool.

So even with a raid setup for redundancy this is doable.

May I ask how you decide to add me content? Do you follow links? Do you use other search engines' results as a starting point?

I could probably shoot many more questions, but don't want to be a nuisance.

Thanks for your time already.

Post reply on HN