Live data from Hacker News

How is search so bad? A case study

svilentodorov.xyz

321–330 of 416 posts

Re: How is search so bad? A case study

#321

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

An AI crawler is needed.

Re: How is search so bad? A case study

#322

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

I've often thought one approach, though one I wouldn't necessarily want to be the standard paradigm, would be exclusivity based on usefulness.

So for example duckduckgo is still trying to use various providers to emulate essentially "early google without the modern google privacy violations", but when I start to think about many of the most successful company-netizens, one thing that stands out is early day exclusivity has a major appeal.

So I imagine a search engine that is only crawling the most useful websites and blogs and works on a whitelist basis. Instead of trying to order search results to push bad results down, just don't include them at all or give them a chance to taint the results. It would have more overhead, and would take a certain amount of time to make sure it was catching non-major websites that are still full of good info ... but once that was done it would probably be one of the best search panes in existence. I have also thought to streamline this, and I know it's cliche, but surely there could be some ml analysis applied to figuring out which sites are SEO gaming or click-baiting regurgitators and weed them out.

Just something I've been mulling over for a while now.

Re: How is search so bad? A case study

#323

Earlier quoted context omitted.

Why the assumption of left or right, when at least in the US most people identify as independents? And to be even more fair, you aren’t talking about “people” here but the OP who can clearly state their own personal preferences, without the need for you to construct a theory of hidden bias.

> Why the assumption of left or right, when at least in the US most people identify as independents? Most people don’t identify as independents, and, anyhow, studies of voting behavior show that most independents have a clear party leaning between the two major parties and are just as consistently attached to the party they lean toward as people who identify with the party. So not only is it the case that most people…

On the contrary, according to relatively recent Pew numbers from 2018 [1], around ~38% of American's identify as independents, with ~%31 democrat and ~%26 republican. To be fair many of them lean towards one party or another (your last point acknowledged), but independents are much more important in politics than they are usually given credit for.

[1] https://www.pewresearch.org/fact-tank/2019/05/15/facts-about...

Re: How is search so bad? A case study

#324
post #222

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

> I can possibly not find anything deep enough about any topic by searching on Google anymore. > It kills my curiosity and intent with fake knowledge and bad experience. I need something better. It's hard for me to take this seriously when wikipedia exists, and almost always ranks very highly in search results for searches for "knowledge topics". Between wikipedia and sources cited on wikipedia, I find the depth of a…

Basically, I think you can divide search between commercial interest search and not commercial interest searches. I can find deep discussions of algorithms curated quite nicely. But information curtains, say, that will be as bad as the OP says.

Re: How is search so bad? A case study

#325
post #289

Earlier quoted context omitted.

I see some claims that PV edited videos, selectively removing portions to give a particular impression, which may well be true. But when CNN did that[1,2], changing "take that shit [violence] to the suburbs" into "a call for peace", the left wing media ignored it. Wikipedia editors didn't add it to CNN's page[3] or add a message that CNN has a history of deceptively editing videos. That's why it's so important to get…

Two wrongs don't make a right – but more importantly, Project Veritas deliberately uses lies and deception to obtain its footage, which it then edits to remove context. That crosses a line far beyond run-of-the-mill reporting bias.

Michael Moore films and edits his documentaries in a similar fashion.

Re: How is search so bad? A case study

#326
post #64

Earlier quoted context omitted.

You could even have all this under one roof: one common search spider that feeds this ensemble of different ranking algorithms to produce a set of indices, and then a search engine front end that round-robins queries out between the different indices. (Don’t like your query? Spin the algorithm wheel! “I’m Feeling Lucky” indeed .)

The Common Crawl is a thing already. Unfortunately, a "full" text crawl of the internets is a YUUUGE amount of data to manage, and I can't think of anything that could change that in the foreseeable future. That's why I think providing a federated Web directory standard, ala ODP/DMOZ except not limited to a single source, would be a far more impactful development.

You don’t really need to store a full text crawl if you’re going to be penalizing or blacklisting all of the ad-filled SEO junk sites. If your algorithm scores the site below a certain threshold then flag it as junk and store only a hash of the page.

Another potentially useful approach is to construct a graph database of all these sites, with links as edges. If one page gets flagged as junk then you can lower the scores of all other pages within its clique [1]. This could potentially cause a cascade of junk-flagging, cleaning large swathes of these undesirable sites from the index.

[1] https://en.wikipedia.org/wiki/Clique_(graph_theory)

Re: How is search so bad? A case study

#327
post #197

Earlier quoted context omitted.

The ‘Google is an advertising company’ is said often on HN. I agree to some extent but, doesn’t that imply that every newspaper company is also just an advertising company? Google solves a real problem and this works well, for them, with an advertising based revenue model. Do they compromise their search to that end? Probably. Do newspapers? Hopefully not, but maybe. To me, that doesn’t make them advertising companie…

> * but, doesn’t that imply that every newspaper company is also just an advertising company?* Historically, at least, they sold subscriptions.

Subscriptions for the “modern newspaper” did not pay the bills, but were proof that people were actually reading the newspaper.

Prior to that there were papers which did indeed make their money from subscriptions. But their content was different as well: explicitly ideological and argumentative. The NYT or Wapo idea of neutral journalism was a later development.

Re: How is search so bad? A case study

#328

Earlier quoted context omitted.

If they were ripe for disruption and it was easy to do this disrupting just be returning better search results, and returning better search results was an easily doable thing then I suppose all the other functioning businesses that have a stake in web search would already be doing that disrupting. Search disrupted catalogs. What will disrupt search?

Boutique hand crafted artisanal catalogs? Not joking I have a feeling subject specific topics will be further distributed based on expertise & trust.

> Boutique hand crafted artisanal catalogs?

I think those are called books. ;-)

Re: How is search so bad? A case study

#329

Earlier quoted context omitted.

PV isn't "gotcha journalism". They've repeatedly committed felonies and completely fabricated things (ex. when they tried to trick WaPo into pushing a fake #metoo story about Roy Moore) in attempts to create their content.

> They've repeatedly committed felonies That's quite a claim. If that's true, why aren't they in jail? I'm sure they have many powerful enemies who would love to see them convicted of these multiple felonies.

https://web.archive.org/web/20100531174024/http://neworleans...

O'Keefe plead out to a misdemeanor and served probation + a fine. But the crime he committed was a felony.

> In January 2012, O'Keefe released a video of associates obtaining a number of ballots for the New Hampshire Primary by using the names of recently deceased voters.

That's probably a felony in most parts of the US, although O'Keefe claims it wasn't since he didn't actually vote.

Re: How is search so bad? A case study

#330
post #262
post #55

Earlier quoted context omitted.

It is reasonable. It is also likely that whatever meta information reddit is sending back (in headers or tags) is probably not dated correctly for the time of the origin post. Google COULD offer more time machine features and perform diffing on pages. But a reddit "page" will always have content changes, as everything is generated from a database and kept fresh on the page. The ONLY metric therefore Google could use…

> It is reasonable. It is also likely that whatever meta information reddit is sending back (in headers or tags) is probably not dated correctly for the time of the origin post. That could explain the first screenshot, but definitely not the second, where google has it tagged as years old.

That's DDG, not google.
Post reply on HN