Live data from Hacker News

Mwmbl: Free, open-source and non-profit search engine

mwmbl.org

111–120 of 129 posts

Re: Mwmbl: Free, open-source and non-profit search engine

#111
post #4

OK, the obvious question: Why go with an unpronounceable name? I mean, great that it was made, but I can't even tell people I'm using... mwumble? But it's spelled em-doubleyou-em-bee-el dot org.

They are begging to be either ignored or forked.

There are no other outcomes if they don't already understand why everyone is telling them this is unusable.

Re: Mwmbl: Free, open-source and non-profit search engine

#112
post #7
post #4

OK, the obvious question: Why go with an unpronounceable name? I mean, great that it was made, but I can't even tell people I'm using... mwumble? But it's spelled em-doubleyou-em-bee-el dot org.

It's pronounced mumble. An explanation is at the very bottom of the github Readme, quoting: > How do you pronounce "mwmbl"? > Like "mumble". I live in Mumbles, which is spelt "Mwmbwls" in Welsh. But the intended meaning is "to mumble", as in "don't search, just mwmbl!"

It's pronounced "google mumble".

Re: Mwmbl: Free, open-source and non-profit search engine

#113

Earlier quoted context omitted.

Primarily a list of options to choose from, preferably that from a non-affiliated site, asking the same in GPT-4 I get the following: >Tangerine Business Savings Account: This account offers a high interest rate of 2.65% to 3.25% on your balance, no monthly fees, no minimum balance requirement, unlimited transactions, free e-transfers, and access to over 3,000 ATMs. >Wise Business Account: This account offers low-cos…

Well if we imagine a search engine as a document retrieval machine, who would publish such a document?

Before I kagi this (not even google, just kagi!) shall we wager on whether there is or is not at least one, likely several such documents? Come on.

Re: Mwmbl: Free, open-source and non-profit search engine

#114

Earlier quoted context omitted.

Primarily a list of options to choose from, preferably that from a non-affiliated site, asking the same in GPT-4 I get the following: >Tangerine Business Savings Account: This account offers a high interest rate of 2.65% to 3.25% on your balance, no monthly fees, no minimum balance requirement, unlimited transactions, free e-transfers, and access to over 3,000 ATMs. >Wise Business Account: This account offers low-cos…

Well if we imagine a search engine as a document retrieval machine, who would publish such a document?

> who would publish such a document?

The banks? In this case. Because if you do the manual searching, you will “manually” go to each bank site, go to accounts, business section and read, a good search engine will do that for me, no middle man (aka some 3rd party sites) and summarize it based on my query, a bad search engine however, will look into a 3rd party website that already created a list, recommended some based on affiliate links, boosted itself in the results by playing the SEO keywords game.

Re: Mwmbl: Free, open-source and non-profit search engine

#115
post #31

Earlier quoted context omitted.

You don't need to remember it, just bookmark and tag however you like (it's anyway a waste of keystrokes to manual type the full domain for such a frequently used site like a search engine)

This is 100% wrong. One's normal phone and laptop is actually only a fraction of uses, and even one's normal device isn't just one thing that needs to be done one time in grade school and then set for life. It's a dozen different things, and they are all perpetually rotating, and most people are not highly optimized with profiles they actually export and import. This idea is great but it's going absolutely nowhere wi…

Name at least half of those perpetually rotating things and quantify "only a fraction"

I'm an actual human, I use alternative search engines, I don't memorize their full names, and the only thing perpertually rotating is the planet

Re: Mwmbl: Free, open-source and non-profit search engine

#116
post #115

Earlier quoted context omitted.

This is 100% wrong. One's normal phone and laptop is actually only a fraction of uses, and even one's normal device isn't just one thing that needs to be done one time in grade school and then set for life. It's a dozen different things, and they are all perpetually rotating, and most people are not highly optimized with profiles they actually export and import. This idea is great but it's going absolutely nowhere wi…

Name at least half of those perpetually rotating things and quantify "only a fraction" I'm an actual human, I use alternative search engines, I don't memorize their full names, and the only thing perpertually rotating is the planet

A stack of old laptops since they are too good to throw away since I buy good stuff and am a Linux user, so even my 10 year old machines are actually still great. So I use them for vacation to take a clean wiped machine, for things like attaching to a 3d printer or being a part of my electronics workbench or out in the garage. Old android tablets and phones which get used about the same way, vacation, device interface. Not to mention, a smaller but similar collection belonging to my wife. These are just the things I might search in a web browser on.

There are actually browsers also built in to 4 TVs, also in the rokus and google TVs attached to those same TVs, also in the Xbox and ps3. But I won't even count any of those. I have actually used them, but I'll give you those for free since I don't actually use those browsers very much.

Also that just reminded me that all of the old devices are fairly regularly getting reinstalled with some new version of a linux or bsd distro fresh every time I pick one back up, so, no configured profiles.

The windows partition on my main machine is frequently reinstalled since I experiment with trying to use either a partition or Frameworks custom usbc module or a regular usbc external drive, or just a partition on a bigger faster external drive. That's one physical device but a few different OS's, and most of those OS's besides my main daily driver get moved around and reinstalled a lot so they are always new and unconfigured., yet, I still need to use them, and that means I use a browser to search from within them.

My kobo, and 3 or so other eink readers. Which, again, occasionally gets reinstalled, so even the one device needs to be set up more than once.

The only reason I don't have to set up a new phone every 6 months is because I value a headphone jack more than most everyone else. So if you would say my usage pattern is an outlier, I would say, 1 so what? Outliers exist and could even be argued to outnumber the center peak of the bell curve, and 2 some of my outlier usage pattern goes the opposite way, like using the same phone for 5 years.

And then of course I use many machines which are not mine. And this is not even counting that my work used to involve some amount of user it support where I would use a users desk or a hot desk at a customer site, I just mean my own personal normal activity is on many other machines besides my own, including relatives, friends, & public machines.

I had to type "google" (back when I used google primarily, and it wasn't already everyone's default) countless times, even though it was the home page on my own main machine.

This question didn't really even deserve the dignity of any answer it is so obtuse.

Re: Mwmbl: Free, open-source and non-profit search engine

#117
post #95

Earlier quoted context omitted.

There was an effort in the early 90s to have search as a protocol so you could have a query and then select the domains you want to run it on and return an aggregate result. It was 100% abandoned and I think that's a mistake. It'd be nice to explore some of those ideas again

You’re thinking of WAIS, I believe: https://en.wikipedia.org/w/index.php?title=Wide_area_informa... >

Yes!

Re: Mwmbl: Free, open-source and non-profit search engine

#118
post #96

Earlier quoted context omitted.

There's more to it than that. What if instead of crawling the php generation of database rows with a bunch of cruft, the administrator published some kind of schema with scraping and querying rules and you could alternatively make a single call to capture all of the data in a sematic schema. You can still do all the stuff you're talking about but it could make search more coherent. An entry for that humans and an ent…

Wasn’t this what the Semantic Web was supposed to enable?

That was more an ontological web. That's a different project which I totally support but this time through ML

Re: Mwmbl: Free, open-source and non-profit search engine

#119
post #115

Earlier quoted context omitted.

Name at least half of those perpetually rotating things and quantify "only a fraction" I'm an actual human, I use alternative search engines, I don't memorize their full names, and the only thing perpertually rotating is the planet

A stack of old laptops since they are too good to throw away since I buy good stuff and am a Linux user, so even my 10 year old machines are actually still great. So I use them for vacation to take a clean wiped machine, for things like attaching to a 3d printer or being a part of my electronics workbench or out in the garage. Old android tablets and phones which get used about the same way, vacation, device interfac…

The only rotating thing in this story is the person actively wasting time erasing all the traces of history without any easily available sync that would prevent the need to type "google" countless times.

But even then compared to that effort remembering a new word is trivial

> So if you would say my usage pattern is an outlier, I would say, 1 so what?

I'd say it's not relevant to this conversation where you barge in with an uber-confident "100% wrong" when it's only "1 person" wrong

Re: Mwmbl: Free, open-source and non-profit search engine

#120

Earlier quoted context omitted.

Most implementations of this have a race towards generalities. The biggest problem used to be when seemingly the whole internet was satisfied with an answer that is extremely wrong and broken when you do it. Chatgpt can work though this without getting into a weird markov cycle maybe half the time which is great. Patterns like "Hey I tried that. It still doesn't work, can you give me another option"

ChatGPT has other failure modes. When a question doesn't have an answer written down somewhere, it really struggles. A case is something like "how do I write a parquet file in Java without using Hadoop". This not at all trivial but quite possible[1], but ChatGPT will in 100% of the time either hallucinate APIs, disregard the instructions to not use Hadoop or give otherwise plausible but incorrect-looking answers. The…

You can call it out "hey you just made that up. Think hard and give me a real answer"

I don't know if "think hard" does anything but it seems to work and if I was the one making chatgpt I'd certainly have configurable keywords like that to tweak the generation settings - mostly so I could skate by on cheaper queries 90+% of the time and then have a fix when they fail

Post reply on HN