Sorry, I can see that sounded bad. I just meant that it sounds like we have a ways to go. I will help!
Not at all, feedback is very welcome!
I wouldn't be surprised if you don't want to talk about it because not could try and avoid the pitfalls, but what is your plan for avoiding bots trying to overwhelm the search scraper bot?
Are you looking to build a trusted network who can verify and validate other users responses on an undefined period?
I use duck duck go. Recently someone showed me their screen where they were using google to do a search. I was absolutely aghast. The last time I used google when you searched for something you saw a simple text list of sites (which is how DDG still works). Instead the google results were… a disaster. You had to scroll through some much garbage before finding actual search results - a list of sites. It was like googl…
Who's going to finally tell them who DDG really relies on?
I know where DDG gets the results from, but in all the years using it, it has never failed to find what I am looking for. Or, I’ve never needed to check google because DDG didn’t find what I was looking for.
Isn't that one more data sets for ML other research purposes instead of a highly up to date search index (For example with news from a few minutes ago).
Roughly speaking, yep - Common Crawl provides a sizable chunk of web data (420 TiB uncompressed, over 3 billion unique URLs, as of May 2022; historic statistics here[1]), and is updated on monthly basis. Not near-real-time, true, albeit relatively fresh. A question to ask could be: how often do users care about information from a few minutes ago, compared to information that has been available for a longer duration o…
I mean, any time someone wants information on current or recent events is your use case right there. If you exclude news entirely, you could maybe disregard recent websites but I imagine that's statistically a pretty large portion of search.
99%+ of people in the medical industry did not sanitize hands or equipment 100 years ago. 99%+ of people in the tech industry currently do not care to do the extra steps required for data neutrality, and privacy. 99%+ of people are lazy to the point of harming themselves and others. 1%- of people examine how the 99%+ do things and pioneer harm reduction tactics in spite of everyone constantly reminding them that no o…
>99%+ of people in the medical industry did not sanitize hands or equipment 100 years ago. this isn't true. 170 years ago, maybe. handwashing became a thing in the late 1800s after Semmelweiss and Pasteur
I will adjust this analogy by 70 years in the future, just for you.
I have no idea what gp means by this. It can be surprisingly hard, but it's not always. Go buy a cooler of water bottles and sell them for 50c in the summer on the side of the street
It seems pretty obvious in an internet context. It's very difficult to make money with a product if the competition is giving their products away for free because it gets paid in a different way. Even your "sell water on a hot day" idea probably won't sell a lot if you set up shop right next to an enormous promo stand from a global bottled water company that gives away bottled water for free. (And due to the magic of…
Ok your explanation makes sense but I don't see how what you said related to the single qualifier presented: being intrinsically valuable, and gp states it as if it were an accepted rule of thumb for all economics, not specific to the internet.
Interesting, can you give an example page that’s not listed?
Sure. A very simple one would be https://ipbl.herrbischoff.com , a public blocklist page referencing a resource used by a couple dozen users. The HTML doesn't get a lot simpler than that and is entirely valid markup. It's (unsurprisingly) low ranked in Google but it's there. Not so in Bing, it's simply not there. The page exists since March 2021. Bing Webmaster Tools reads "Discovered but not crawled. URL cannot appe…
This is fascinating, thanks. Have you experimented on allowing the bing ad bot that you have blocked? If they have some kind of retaliatory non-crawling?
I have no idea what gp means by this. It can be surprisingly hard, but it's not always. Go buy a cooler of water bottles and sell them for 50c in the summer on the side of the street
It seems pretty obvious in an internet context. It's very difficult to make money with a product if the competition is giving their products away for free because it gets paid in a different way. Even your "sell water on a hot day" idea probably won't sell a lot if you set up shop right next to an enormous promo stand from a global bottled water company that gives away bottled water for free. (And due to the magic of…
No, even then it is the same concept. The big global bottled company thought that the intrinsic value of the water is more than 50c or maybe more than 50$ . Now you have bunch of people carrying their water bottles around, giving them even more advertising. People think that water bottle company is what I should get based on those carrying it around. So infact, they are giving away water bottle for free, which compared to competitors might be 50c or 50$. but they think it is worth losing that for ad
The steps I took were: 1. Go to kagi.com 2. Right-click on the search field and choose "Add a keyword for this search..." 3. Fill in the keyword (e.g. kagi, as I did) 4. Right click in the top url+search field, and choose "Add Kagi search" Now you can search via keyword, you can change to kagi while typing, and you can set the default search engine to kagi in the preferences.
Thanks. Apparently, it works fine for Kagi but not for mwmbl.org (from the article.) I foolishly didn't try a different search engine :P
That is odd! The search bar of mwmbl has the same properties as kagi's. Perhaps you're right, and Mozilla does limit which sites get the treatment.
>To your last point, I don’t quite agree. Google’s incentives are misaligned such that keeping you on Google.com just a little longer is better than not because you are more likely to click on an ad. I agree, which leads me to the conclusion that subscription is the best way to avoid this conflict of interest. Unfortunately, most of the world won't subscribe to a search engine, and doesn't seem to mind ads - to a deg…
Here's a search engine that I'm subscribed to: https://kagi.com In the 20-30 searches that I do in a day, I still have to google about half of them. Either because it's stuff Google does well (currency conversion, for example), or Kagi just doesn't get what I'm trying to search. I remember starting out with the Internet searching on Altavista and Yahoo and Lycos. The information that was present was nowhere near as n…
Currency conversion is nothing you have to sell yourself to Google for. Just bookmark a bank, a financial or an academic research site that seems trustworthy. I have used the same ones for over 20 years, probably found them using Altavista at the time...
How should I pronounce this search engine? I know naming is hard, but if you want something to be easily adopted, having a sticky and pronounceable name is paramount!