Live data from Hacker News

More content by people, for people in Search

blog.google

311–320 of 320 posts

Re: More content by people, for people in Search

#311
post #307
post #106

Earlier quoted context omitted.

Same here. For more niche stuff, + "site:news.ycombinator.com" is standard practice. Or Algolia's HN search. Another one is + "forum" or + . I train myself to use DDG, but, to my own dissapointment, I end up with "!g site:something-trustworthy.com" way too often. For regular Google searches, I've had the feeling of being "scammed" for years, though. The "fishing devil" on the no-results page also feels like buttering…

Here is a what you should add to search request to find mostly old style forum discussions: inurl:forum|viewthread|showthread|viewtopic|showtopic|comments|"index.php?topic" | intext:"reading this topic"|"next thread"|"next topic"|"send private message"

This is a really useful tip, many thanks. It has probably largely been a personal comfort zone, but I haven't dug much deeper than site: or ext: operators in recent years. A lazy habit. Also, strangely, I have always had doubt as to the real efficiency of the OR operator in web search (maybe it has yielded unsatisfactory results in some distant past, and for different engines, not only Google; hence the prejudice). I didn't even know that the "pipe" symbol should actually work for that.

Your example is a really good one, because it illustrates the power of OR remarkably well. The key is using it to combine several different operators, not just duplicating a single one, e.g. "query site:a.com | site:b.com". That has been my main way of digging deeper over the years, along with "#" to specify a date range.

Interestingly, in case of "query site:firstsite.ee|secondsite.ee" the "|" doesn't seem to work, I get zero results. Why is that? To narrow down the results based on both the domain name and top level domain, I have to add the inurl: operator. E.g. "query inurl:firstsite|secondsite site:.ee"

Strangely, "piping" is the essential thing I do on the Unix command line every day. I haven't put (that) much thought into using Google's "|" operator in a similar vein, that is, to combine several (3-4) different operators. Possibly partly because one used to get pages and pages of results even with very simple, single-operator queries.

Time to refresh my memory about Google's search operators, I guess. Thanks again for sharing this example.

Re: More content by people, for people in Search

#312

>> We know people don’t find content helpful if it seems like it was designed to attract clicks rather than inform readers. So... Back when "SEO" was new, I would read the Matt Cutts blog. He was head of "anti-spam." I remember thinking back then that anti-spam was an ignorant frame. Once Google gained importance, websites started trying to improve their rankings. That might mean migrating from Flash to HTML. It migh…

I agree with you, but Google also faces a perhaps impossible problem of how to detect good content on YouTube. They can look at how long people watch; but that ends up rewarding rambling pointless blather when a concise two minute video would be perfect. They can look at likes or subscribes, but that causes many wasted human lifetimes per day, of people saying and listening to pleading to click said buttons.

>> Google also faces a perhaps impossible problem of how to detect good content on YouTube

Perhaps, but only because Google is being small minded. Define the "problem" in a certain way, and come up with certain solutions. A broader minded frame would define the problem(s) more broadly. Presenting users with UI. Creating a good, or at least respectful, commercial incentives framework for creators. The litmus for this is the content. What content gets created. Mitigating or avoiding bad consequences, like spiraling towards pointless blather.

Yes, this would require subjective decision making, but at least it doesn't require dim witted, corporate delusion. Call the dog by its name.

Re: More content by people, for people in Search

#313
post #252

Earlier quoted context omitted.

The dissonance is between Google's native perception and external reality: reality 1 : The www exists. Google indexes it, analyzes it and delivers it to users. Users like certain things, like original content. reality 2 : Google's ranking policies/algorithms influence the web. The "original content" that exists in a world without Google is different to the content that exists in a world where Google ranks such web pa…

I think it’s more that nobody knows the policy. It’s all ML changes few people understand or can even communicate. I don’t expect google would explain if they knew how it worked, but they also cannot explain how their technology works.

IDK. I suspect this is less true than most acknowledge.

Black box or not, google are creating these systems, analyzing, optimising and implementing them. It's indeed hard or impossible to trace back individual examples and their whys. But the macro effects are visible and choices are made.

It's just easier to "blame the computer," internally or externally. It's just like bank employees merge always blame "regulations." In fact, whatever they are blaming is a policy created to implement that bank's compliance framework which exists to satisfy the regulator's goals and the bank's goals." Most of the time, the underlying regulation is distantly removed from whatever piece of bureaucracy is annoying the customer or employee. But, if it can be blamed on the regulator, it will be.

It's just easier, if at all possible, to blame an immovable and mysterious force rather than the more likely culprit: humans doing a bad job.

Re: More content by people, for people in Search

#314
post #197

Earlier quoted context omitted.

It’s safe to assume that a social ranking system covering the whole web will be gamed.

There is an extent they can go to with account verification where gaming is not very feasible. Can you create a fake account with a verified credit card number, verified phone number, passport, drivers license, account history consistent with human usage, Google One subscription, etc.? You probably can, but doing it at-scale is going to be quite costly.

Can you get a lot of people to install a piece of software and then use their account instead? Could you even pay those people per hour of usage of their account?

Quite cheaply

Re: More content by people, for people in Search

#315
post #120

Earlier quoted context omitted.

Or, more likely, they are just SEO optimising by providing a lot of prose (which Google has historically prioritised in search rankings) - the same way as all those AI-generated article summaries on the astroturfing sites

How do you propose anyone (Google or humans) differentiate between 500 mostly-similar recipes for brownies? I think HNers look at food blogs with a very depressing robotic expectation.

This is what algorithms like PageRank were originally designed to do. Once upon a time the popular brownie recipes are the ones that everyone linked to, or the ones on the cooking sites that everyone linked to.

To be clear, I don't really care if the person wants to put their life story on their recipe site. I am annoyed that they are forced to do so in a really mechanical and disingenuous way just to make a Google algorithm happy - and as a result I lean heavily on the recipe sites that have alternate marketing channels/revenue streams (i.e. Serious Eats, ChefSteps, etc)

Re: More content by people, for people in Search

#316
post #31

Earlier quoted context omitted.

And with the possibility of having the measures in metric. Otherwise, it is almost useless for most of the world.

I like how you said "possibility", as if 3/8th of a cup wasn't a good enough measurement.

No, it is not a good measurement. I have cups of many sizes.

Re: More content by people, for people in Search

#317
post #90

Earlier quoted context omitted.

Maybe the bootleg page could be a source of inspiration? It has a garbage/content ratio of 32 (the browser downloads 32 bytes for every byte of content) while the original page has a 400 ratio (the browser downloads 400 bytes for every byte of content). It's borderline denial of service attack against the visitor. The original also has a slow aggressive cookie box that is unnecessary as visitors can be spied on using…

> Also if google starts ranking such pages up they may expose themselves to class action lawsuits as users could ask for a refund for the power, hardware and telecom bills incurred in being lead to load them? Sorry, but this is completely insane. Search engines should most certainly and definitely not ever be liable for linking to original sources for such a stupid reason. This American pro-litigious attitude has to…

Sure, it's tongue in cheek, but there is a tragedy of the commons in there: whatever massive waste devs and their employers inflict on society goes unpunished because there is no incentive to change that behaviour. Was suggesting some activist subvert computer misuse laws to fix this. Primary legislation or big tech using its weight would do as well.

In the physical world one can't dump a rusty aircraft carrier on someone else's lawn and get away with it. Should be the same in tech.

Re: More content by people, for people in Search

#318
post #103

Reading through the comments here (and very much feeling the pain they describe), this idea came: Shouldn't we have a search engine that heavily favours the types of websites that we typically look for? You know, the classic 90s style tech blogs, the plain HTML documentation pages. Ignoring websites with ads, sites with lots of baggage (fonts, scripts, whatnot), sites with lots of images. Maybe increase the ranking o…

Are you talking about Marginalia? [https://search.marginalia.nu/]

Re: More content by people, for people in Search

#319
post #90

Earlier quoted context omitted.

> Also if google starts ranking such pages up they may expose themselves to class action lawsuits as users could ask for a refund for the power, hardware and telecom bills incurred in being lead to load them? Sorry, but this is completely insane. Search engines should most certainly and definitely not ever be liable for linking to original sources for such a stupid reason. This American pro-litigious attitude has to…

Sure, it's tongue in cheek, but there is a tragedy of the commons in there: whatever massive waste devs and their employers inflict on society goes unpunished because there is no incentive to change that behaviour. Was suggesting some activist subvert computer misuse laws to fix this. Primary legislation or big tech using its weight would do as well. In the physical world one can't dump a rusty aircraft carrier on so…

> In the physical world one can't dump a rusty aircraft carrier on someone else's lawn and get away with it. Should be the same in tech.

Google's search results are most definitely not your lawn, nor are heavy websites very much alike to rusty aircraft carriers. Or however this analogy is supposed to work anyway.

Re: More content by people, for people in Search

#320
post #311
post #307

Earlier quoted context omitted.

Here is a what you should add to search request to find mostly old style forum discussions: inurl:forum|viewthread|showthread|viewtopic|showtopic|comments|"index.php?topic" | intext:"reading this topic"|"next thread"|"next topic"|"send private message"

This is a really useful tip, many thanks. It has probably largely been a personal comfort zone, but I haven't dug much deeper than site: or ext: operators in recent years. A lazy habit. Also, strangely, I have always had doubt as to the real efficiency of the OR operator in web search (maybe it has yielded unsatisfactory results in some distant past, and for different engines, not only Google; hence the prejudice). I…

Google is not very consistent and this feature is an unsupported vestige of the past.
Post reply on HN