Also maybe hackernews should do a monthly post your blog/website just like the monthly job posts.
We should promote more personal indexing, rather than algorithmic indexing
11–20 of 23 posts
Re: We should promote more personal indexing, rather than algorithmic indexing
#12Comparing the atrocious search landscape of 2023 to personal indexing is to compare poorly.
Instead: those who lived through it should compare personal indexing to the golden years of altavista(.digital.com) and the extremely powerful and unpolluted results it produced and modified with boolean operators.
I never once thought that there was some utility left to be mined from the Yahoo approach to things once I switched to Altavista. I think we're only pining for personal indexing in comparison to the garbage that is Google in 2023.
Re: We should promote more personal indexing, rather than algorithmic indexing
#13https://payperrun.com/%3E/search?displayParams={%22q%22:%22c...
(btw, I just launched this llm-embedding based search service that lets you check if a startup idea has already been tried/failed).
I don't know if this idea has a higher death rate than the baseline, but my guess is Google/PageRank is good enough for most use-cases, and then if you want quality sources, you can just follow them on YouTube, Twitter, Instagram, etc. Wait, maybe I shouldn't try to compete with Google?
Re: We should promote more personal indexing, rather than algorithmic indexing
#14But I think you have a good point.
Re: We should promote more personal indexing, rather than algorithmic indexing
#15Re: We should promote more personal indexing, rather than algorithmic indexing
#16There is the question of what my judgement is and the question of what your judgement is. The primary selection process (that shows me articles) is great right now, I am upvoting maybe 220 articles out of 300 in a cycle. I pick out links to post to HN and I am currently adding links to the queue faster than I am posting them which means I can definitely raise the quality of what I post but the definition of "quality" is where I get stuck. It has all sorts of factors such as a lack of annoyingness (I hate those cookie popups but there is a lot of good news behind them) but there are also articles that look really good to me at first (I like what they set out do) but then what I look at them again I realize they didn't accomplish what they set out do.
I do think votes and comments are worth something, but I also know that I could get more of both by posting clickbait articles. On one level I want to post things that are enlightening, boy I get frustrated that y'all just don't care about robotics or chemical recycling of polymers or Arduino projects. (Though my real secret ambition is to get a #1 post about sports...)
Somehow I want to pose the problem of posting to HN, Mastodon, etc. as a sequential recommendation problem which means I have to back and look at all the papers on the subject that YOShInOn has collected for me. Also I am likely to put some more work into "quality models", particularly a stacked model for predicting votes if not comments on HN articles, a broad topic model based on data from Tildes (is is sports? music? science?) and particularly sentiment models.
That last one is on my mind because I'm thinking about the emotional tone of what I post to Mastodon, some days I think I should just stick to posting flower photos because they get good engagement, but past the people who are calling everybody a "fascist" that get amplified there is a "silent majority" of people on Mastodon who try to avoid the news and other inflaming topics so I am torn between being unrelentingly positive or trying to balance out positive and negative articles to make a more appealing feed to the good people of Mastodon. It would probably be 2-3 days of labeling work to make a sentiment model but if I had to find and categorize 5000 angry toots it would kill me, but I am thinking now about grading my own posts (what the system is going to do inference on anyway) and also grading high-engagement and predicted high-engagement submissions to HN to make a model that finds high-engagement posts that aren't clickbait.
Re: We should promote more personal indexing, rather than algorithmic indexing
#17I'm open minded but ... Comparing the atrocious search landscape of 2023 to personal indexing is to compare poorly. Instead: those who lived through it should compare personal indexing to the golden years of altavista(.digital.com) and the extremely powerful and unpolluted results it produced and modified with boolean operators . I never once thought that there was some utility left to be mined from the Yahoo approac…
A huge amount of the crap on the web is there because of Google. If you think what makes it to the top 10 is bad, you should see the crap that doesn't make it. A different search engine, with different intentions, would have to filter all that out.
If you don't believe me, try setting "showdead" on your HN profile and see all the posts the system makes [dead] before you even see them because people have some crap blog and they post nothing but links to their crap blog. I almost feel bad for those people, had they done that in 2013 they might have actually shown up in Google, today they can do that until they are blue in the face and get about zero traffic, not because Google is effective at filtering that crap out, but because they will be buried underneath all the people doing the same.
A better resource today would have to start with some radical choice such as whitelisting, if only to reduce the head-end costs of ingesting material.
It's tempting to imagine some rules like: no ads, no popups of any kind, government mandated or not, especially no cookie banners, no paywall, but even sites like Wikipedia fail at those criteria today.
Re: We should promote more personal indexing, rather than algorithmic indexing
#18I'm open minded but ... Comparing the atrocious search landscape of 2023 to personal indexing is to compare poorly. Instead: those who lived through it should compare personal indexing to the golden years of altavista(.digital.com) and the extremely powerful and unpolluted results it produced and modified with boolean operators . I never once thought that there was some utility left to be mined from the Yahoo approac…
It's not just Google, it's the web that Google made. A huge amount of the crap on the web is there because of Google. If you think what makes it to the top 10 is bad, you should see the crap that doesn't make it. A different search engine, with different intentions, would have to filter all that out. If you don't believe me, try setting "showdead" on your HN profile and see all the posts the system makes [dead] befor…
> It's tempting to imagine some rules like: no ads, no popups of any kind, government mandated or not, especially no cookie banners, no paywall, but even sites like Wikipedia fail at those criteria today.
This sounds like the approach that the Marginalia (https://search.marginalia.nu/) search engine is taking. My understanding is that its algorithm favors text-heavy sites. And additions to its index are done via GitHub Pull Request so it's effectively using an approve-list (whitelist).
Re: We should promote more personal indexing, rather than algorithmic indexing
#19Earlier quoted context omitted.
It's not just Google, it's the web that Google made. A huge amount of the crap on the web is there because of Google. If you think what makes it to the top 10 is bad, you should see the crap that doesn't make it. A different search engine, with different intentions, would have to filter all that out. If you don't believe me, try setting "showdead" on your HN profile and see all the posts the system makes [dead] befor…
> A better resource today would have to start with some radical choice such as whitelisting, if only to reduce the head-end costs of ingesting material. > It's tempting to imagine some rules like: no ads, no popups of any kind, government mandated or not, especially no cookie banners, no paywall, but even sites like Wikipedia fail at those criteria today. This sounds like the approach that the Marginalia ( https://se…
Re: We should promote more personal indexing, rather than algorithmic indexing
#20The problem isn't black and white. Discussion adds value, if you didn't think so you wouldn't have posted here. so if not black and white then your talking about subtle and degree which is hard for a machine to determine. Also, most people engage with narratives/stories not facts. I'm trying to downplay your comments, just its hard to start a social media company. there's a reason it's considered a tarpit idea.