Live data from Hacker News

A search engine that favors text-heavy sites and punishes modern web design

search.marginalia.nu

621–630 of 735 posts

Re: A search engine that favors text-heavy sites and punishes modern web design

#621
post #558

I searched for "giraffe evolution" (without quotes) and received the following links on the first page: - Evolutionist scientists say the theory is unscientific and worthless - Seven Mysteries of Evolution - OTHER EVIDENCE AGAINST EVOLUTION - Evolution Falsified Not a single result about the evolution of giraffes...

Unfortunately, as you've discovered, giraffes are often used by crackpots to try to disprove evolution. Google seems to get around this by heavily boosting known authoritative sources like National Geographic and NIH. But, sadly, those are JS/image heavy sites.

Re: A search engine that favors text-heavy sites and punishes modern web design

#622
post #412
post #410

Earlier quoted context omitted.

This isn’t entirely Google’s fault. Recipes on their own aren’t copyrighted in the US, and adding this fluff text is a way around that.

Yes it is. They're giving a lower quality result to the user (their customer ... but not for long if competitors can get just a little bit better)

A small nit: I think Google’s customers are actually the companies paying for ads. Its an important distinction and probably explains why their search results’ quality has gone down

Re: A search engine that favors text-heavy sites and punishes modern web design

#623

Earlier quoted context omitted.

Try a query on the form site:www.example.com ;-)

> site:www.washingtonpost.com > Blacklisted false > site:www.wsj.com > Blacklisted false > site:www.rt.com > Blacklisted false > site:www.nytimes.com > Blacklisted true ?

Hmm, not sure what caused it to end up there, but I removed it from the blacklist. It still doesn't seem to want to index the domain however, probably CDN-related.

Re: A search engine that favors text-heavy sites and punishes modern web design

#624

Earlier quoted context omitted.

Vanilla HTTP should be fine. I think 80% of the urls are HTTP. If you're getting no such domain, it's either blocked because it looks too much like a spam domain, or it simply hasn't been discovered yet. What's the TLD? I severely restrict some cheaper TLDs because they gave so much spam. For example, cr.yp.to is an example of a baby I know I've definitely thrown out with the bathwater.

www.ft.com gets 'no such domain'

I added it now, but it turns out it's behind a CDN so I still can't crawl it.

Re: A search engine that favors text-heavy sites and punishes modern web design

#627
post #518

Fantastic idea and it works quite well for short phrases that I tried. As expected I am getting a lot of early 2000s sites which is something that I miss on regular Google. Hilariously searching for "array data structure" got me one of the top results this little tiny page: http://infolab.stanford.edu/~backrub/google.html

> We have designed Google to be scalable in the near term to a goal of 100 million web pages

Funny, that's about where I see my search engine capping out as well.

Re: A search engine that favors text-heavy sites and punishes modern web design

#628

Question: how do we benchmark search engines? Are there any groups attempting to provide (open) solutions in this space? (It seems to me that if you want to build a good search engine, this is the question you need to address first.)

The search term you might be looking for is "information retrieval" there are pretty standard measurements for whether you are getting good results, but they are generally conditioned on stuff like click through rate, comparing to expert ranking and other signals that the user gives you that it was a good or bad return of search results.

Re: A search engine that favors text-heavy sites and punishes modern web design

#629
post #340

Earlier quoted context omitted.

This sort of optimization is why simple recipes are typically found at the end of a rambling pointless blog post now. Still, the best way to break SEO is to have actual competition in the search space. As long as SEO remains focused on Google there is an opportunity for these companies to thrive by evading SEO braindamage.

It's also because that's a way of trying to copyright protect recipes, which are normally not copyright protected. > “Mere listings of ingredients as in recipes, formulas, compounds, or prescriptions are not subject to copyright protection. However, when a recipe or formula is accompanied by substantial literary expression in the form of an explanation or directions, or when there is a combination of recipes, as in a…

But that copyright protection only extends to the literary expression. The recipe itself is still not covered by copyright, even if accompanied by an essay.
Post reply on HN