Live data from Hacker News

Blocked Sites is discontinued

support.google.com

121–130 of 141 posts

Re: Blocked Sites is discontinued

#121
post #64
post #59

Earlier quoted context omitted.

Sure it is, obviously you don't take any small amount of blocks as a signal. By getting a significant amount though and looking at each persons block list to normalize there blocks I think it would be useful.

> Sure it is, obviously you don't take any small amount of blocks as a signal. You do realize spammers are extremely good generating large amounts of things, right? > By getting a significant amount though and looking at each persons block list to normalize there blocks I think it would be useful. Spammers are also very good (though not great) at adding innocuous data alongside their spam to look like real users. You…

The same thing is true of backlinks. Sure there is a point in which it may not be worth the trouble, not questioning that.

Re: Blocked Sites is discontinued

#122

Earlier quoted context omitted.

I'm not sure google is even clear about what its core business is, these days. It really seems quite clear that they've abandoned much of the approach and focus on being the information finder to now being the social-network wannabe. It's as if something flipped 180 degrees - they went from being the company that wanted to help you find out about everything to being the company that wanted to find everything about yo…

Google's core business is advertising.

Anyone who disagrees with this other Jacques is invited to take a squiz at Google's annual reports.

Any of them since the introduction of adwords and adsense. It doesn't matter which. They all basically read the same.

Re: Blocked Sites is discontinued

#123

Earlier quoted context omitted.

Works great for a company that risks generating a lot of noise to consumers (i.e. producing and marketing physical products). But for Google it's pervasiveness and breadth of services were part of the appeal, it's not like it's minor products interfered with the major parts of the company nor cost them much to maintain.

I agree that Jobs' advice was a mismatch for Google. Trying a million things and seeing what sticks, like an incubator, is a better idea for Google. The problem is when that turns into Microsoft-esque trend chasing. But that only happens when upper management focuses on trend chasing. Otherwise they just end up as passing experiments.

I call it the Spaghetti Cannon Strategy.

Load up a cannon full of ideas, fire and see what sticks.

Microsoft were doing it before Google, but the financial-strategic dynamics of those businesses is very similar.

There's a fountain of cash and a series of spaghetti cannons around it blasting away furiously. Almost nothing sticks. But that's OK, because there's just so much money.

Re: Blocked Sites is discontinued

#124
post #115

Earlier quoted context omitted.

They seem to be showing up for me. I don't have any recent organic examples, but here's query for one of the most popular threads, and the quora thread is the first result: https://www.google.com/search?q=How+do+I+get+over+my+bad+hab... What do you mean by "private site"? The other objectionable/EE-like thing about it is how it hides the answers unless you log in. That really sucks.

EE never hid their answers. They were only obfuscated. First[1] they made it look like the page ended, but you could scroll further and get the answers. Later, I saw that they had a section that looks like the answers, but all of the text was blurred. Again, if you scrolled past the 'end of the site' you could still get to the answers. Looks like Quora gives you the question, and the first answer, but all other answe…

If you add ?share=1 to the end of a Quora URL you can get the full content.

Re: Blocked Sites is discontinued

#125

I find both its birth and death interesting. The birth because the fact that you could mark sites as spam was one of the early talking points at Blekko, the back end architecture of our engine includes a 'selector' mechanism (slashtags) and maintaining a personal 'spam' slash tag came along for free. Its one of the features I continue to use. And then it showed up as a feature in Google's results which I found intere…

If google is still researching AI then it is likely because the war with the spammers can't be won in any other way. Sooner or later you end up in a situation where the amount of 'spam' versus the amount of 'ham' is such that no matter how good your algorithms and how good your computing infrastructure you'll end up spitting out a lot of spam. Human curation is a stop-gap solution, computers can generate spam faster…

It would indeed be an interesting situation if the spammers forced the creation of AI (either they doing it to make better spam or someone like Google to deny it).

One of the interesting things I have experienced in my time at Blekko has been that "growth" on the Web isn't really growing all that much. Sure there are trillions of pages being created but there are only so many things the few billion people in and around the Internet care about. There are 'hard information' places, which are things like libraries where reference searches are common, there are 'entity' places, be thay shops or service providers or SOMA startups, and there are "transient" places where information is current and then stale, to be stored and later reconstructed like coral into a 'dead' (in terms of change) but 'useful' base. Seeing the web from the point of view of a web scale crawler and indexer it starts to be clear that the mantra "Organize all the world's information" is getting tantalizingly close to a dynamically stable froth.

I have to believe that Google has figured this out, some of the smartest engineers I've worked with are at Blekko but Google has its share as well. When you trawl through the fishery and all you get are trash fish you start to wonder, "hmm did we actually catch all the fish there are?" So to it goes with "the Web".

I started doing some speculation [1] on how you could value information that was discoverable on the Internet. And one of the schemes I came up with is how many people would find that information "useful", where useful really means they would have some reason of seeking it out. And then scaling that value by the value to them of having it. So for example if my genome was online, there is maybe a dozen people who would find it "useful", and of that dozen probably on the insurance actuaries and perhaps the occasional researcher who would find it "valuable".

Now one takes that unit of applicability/value and scales it again by the "cost" to acquire it (find it, index it, etc). And from that you can compute the total size of the Web. Well estimate it at least. So far my upper bound (based primarily on the fact that world population is stabilizing) is about 72 trillion documents at any given time.

When you look at it that way you can see that ultimately the spammers lose. They lose because over time the actual information that rises to the useful vs cost threshold is identified and classified, or the legitimate channels that provide dynamic information, or the legitimate archival sources that provide distilled information are all, for the most part known and 99.99% of all your users can find everything they want. And as a spammer you are no longer given the free reign of "appear and be indexed" you have to ask for admittance though some form or another. And the level of new credible sources that are created is inherently a function of the number of people that exist, and the number of people that exist is stabilizing.

When Yahoo started with its human curation it was vastly better than anything anyone else could do, and then it was overwhelmed by a combination of growth and algorithms that could much more rapidly infer curation from the social signals of bookmark pages and article reference. Curation has come back into favor, and it combined with machine learning algorithms will create what is essentially a stable corpus of documents[2] known as "the web".

Wikipedia is a great analog for what is happening world wide. All the reference articles they want to put in are nearly done, the number of editors required has gone down not up, and the future of web search is, in my opinion of course, similar. The only new ground on the web is social networking and Google understands that it seems. Its fortunate for them that only deep pocketed entity that is possibly a near time threat there is Microsoft, and since to date Microsoft is trying to do exactly what they did, they benefit from having already been through that part and know exactly what Microsoft will have to do next.

[1] My side hobby is attempting to discern the economics of information.

[2] Documents are just that, pieces of information, I hardly count every page rendered by Angry Birds in a browser as a separate document, in fact applications are them selves a single "document" in the since that some number of people will seek them out, and connect/consume them over what we think of as the "web" interface.

Re: Blocked Sites is discontinued

#126

Earlier quoted context omitted.

Mine was surprisingly short.. www.w3schools.com

Well I had all those sites that scrape StackOverflow added to my list. I love that SO uses a CC license but it does have some unintended consequences.

Is there a public list of such site? I think I'd find it quite useful (for my own blocking).

Re: Blocked Sites is discontinued

#127

Earlier quoted context omitted.

I agree that Jobs' advice was a mismatch for Google. Trying a million things and seeing what sticks, like an incubator, is a better idea for Google. The problem is when that turns into Microsoft-esque trend chasing. But that only happens when upper management focuses on trend chasing. Otherwise they just end up as passing experiments.

I call it the Spaghetti Cannon Strategy. Load up a cannon full of ideas, fire and see what sticks. Microsoft were doing it before Google, but the financial-strategic dynamics of those businesses is very similar. There's a fountain of cash and a series of spaghetti cannons around it blasting away furiously. Almost nothing sticks. But that's OK, because there's just so much money .

Microsoft also fell into the trend-chasing trap before Google. It's a very fine line, though. Excel was probably trend-chasing at its time, but they were doing so many things at once back then that going head-to-head with Lotus 1-2-3 wasn't a bad idea to throw in there. Whereas stuff like Windows Phone or Zune seemed like major strategic shifts dictated from the top, not just "oh let's make one of those too".

Re: Blocked Sites is discontinued

#128

Earlier quoted context omitted.

Works great for a company that risks generating a lot of noise to consumers (i.e. producing and marketing physical products). But for Google it's pervasiveness and breadth of services were part of the appeal, it's not like it's minor products interfered with the major parts of the company nor cost them much to maintain.

I agree that Jobs' advice was a mismatch for Google. Trying a million things and seeing what sticks, like an incubator, is a better idea for Google. The problem is when that turns into Microsoft-esque trend chasing. But that only happens when upper management focuses on trend chasing. Otherwise they just end up as passing experiments.

The corollary of the spaghetti strategy is that if something doesn't stick, you kill it. Taking Jobs' advice is really just setting a higher bar for sticking.

Re: Blocked Sites is discontinued

#129
post #33

Earlier quoted context omitted.

>Their offered download link doesn't even work for me, it just notifies me of its shutdown. After you go to that link and read the notification, there is another link to download the blocklist as a plain-text file. Mine contained a number of domains I've added during the years. It's quite interesting to see all the domains I've blocked so far :)

Mine was surprisingly short.. www.w3schools.com

They also use these domains: w3schools.com ww.w3schools.com wwww.w3schools.com

Probably others too.

Post reply on HN