Live data from Hacker News

Google de-indexed Bear Blog and I don't know why

journal.james-zhan.com

61–70 of 199 posts

Re: Google de-indexed Bear Blog and I don't know why

#61
post #35

Earlier quoted context omitted.

For a lot of things we don’t opt for the cheapest solutions that also lack redundancy for a lot of things. Why not for the “information highway”? Most efficient = cheaper. A lot of times cheaper sacrifices quality, and sometimes safety.

It's not even that. It's that "centralization is more efficient" is a big fat lie. If you look at the "centralized systems" they're... not actually technologically centralized, they're really just a monopolist that internally implements a distributed system. How do you think Google or Cloudflare actually work? One big server in San Francisco that runs the whole world, or lots of servers distributed all over?

I know exactly how they work, but they have a single entry point, as a customer you don't really care that the system is global, and they also have a single control plane, etc. Decisions are efficient if they need to be taken only once. The underlying architecture is irrelevant for the end user.

Why do you think they're a monopoly in the first place? Obviously because they were more efficient than the competition and network effects took care of the rest. Having to make choices is a cost for the consumer - IOW consumers are lazy - so winners have staying power, too. It's a perfect storm for a winner-takes-all centralization since a good centralized service is the most efficient utility-wise ('I know I'm getting what I need') and decision-cost-wise ('I don't need to search for alternatives') for consumers until it switches to rent seeking, which is where the anti-monopoly laws should kick in.

Re: Google de-indexed Bear Blog and I don't know why

#62
post #13
post #4

We need a P2P internet. No more Google. No more websites. A distributed swarm of ephemeral signed posts. Shared, rebroadcasted. When you find someone like James and you like them, you follow them. Your local algorithm then prioritizes finding new content from them. You bookmark their author signature. Like RSS but better. Fully distributed. Your own local interest graph, but also the power of your peers' interest gra…

No, you need to bust up Google as the monopolist it is. YouTube should get split out and then broken up. Google Search should get split out and broken up. etc. This is not a problem you solve with code. This is a problem you solve with law.

> This is not a problem you solve with code. This is a problem you solve with law.

When the DMCA was a bill, people were saying that the anti-circumvention provision was going to be used to monopolize playback devices. They were ignored, it was passed, and now it's being used to monopolize not just playback devices but also phones.

Here's the test for "can you rely on the government here": Have they repealed it yet? The answer is still no, so how can you expect them to do something about it when they're still actively making it worse?

Now try to imagine the world where the Free Software Foundation never existed, Berkeley never released the source code to BSD and Netscape was bought by Oracle instead of being forked into Firefox. As if the code doesn't matter.

Re: Google de-indexed Bear Blog and I don't know why

#63
post #20
post #13

Earlier quoted context omitted.

No, you need to bust up Google as the monopolist it is. YouTube should get split out and then broken up. Google Search should get split out and broken up. etc. This is not a problem you solve with code. This is a problem you solve with law.

Yes. It's a political problem and a very old one. That's why we also already have solutions for it, antitrust laws and other regulations to ensure competition and fairness in the market, to keep it free. Governments just have to keep funding and enabling these institutions.

There are very few things I'd consider a silver bullet for a lot of problems, but antitrust enforcement to break up near-monopolies is one of them.

Re: Google de-indexed Bear Blog and I don't know why

#64
post #50

Earlier quoted context omitted.

So the spammer would link to my search page with their query param: example.com/search?q=text+scam.com+text On my website, I'll display "text scam.com text - search result" now google will see that link in my h1 tag and page title and say i am probably promoting scams. Also, the reason this appeared suddenly is because I added support for unicode in search. Before that, the page would fail if you added unicode. So th…

Interesting - surely you'd have to trick Google into visiting the /search? url in order to get it indexed? I wonder if them listing all these URLs somewhere are requesting that page be crawled is enough. Since these are very low quality results surely one of Google's 10000 engineers can tweak this away.

> surely you'd have to trick Google into visiting the /search? url in order to get it indexed

That's trivially easy. Imagine a spammer creating some random page which links to your website with that made up query parameter. Once Google indexes their page and sees the link to your page, Google's search console complains to you as the victim that this page doesn't exist. You as in the victim have no insight into where Google even found that non-existent path.

> Since these are very low quality results surely one of Google's 10000 engineers can tweak this away.

You're assuming there's still people at Google who are tasked with improving actual search results and not just the AI overview at the top. I have my doubts Google still has such people.

Re: Google de-indexed Bear Blog and I don't know why

#65
post #58

When I reload the page " https://journal.james-zhan.com/google-de-indexed-my-entire-b... ", I get Request URL: https://journal.james-zhan.com/google-de-indexed-my-entire-b... Request Method: GET Status Code: 304 Not Modified So maybe it's the status code? Shouldn't that page return a 200 ok? When I go to blog.james..., I first get a 301 moved permanently, and then journal.james... loads, but it returns a 304 not modi…

You get a 304 because your browser tells the server what it has cached, and the server says "nothing changed, use that". In browsers you can bypass the cache by using Ctrl-F5, or in the developer tools you can usually disable caching while they're open. Doing so shows that the server is doing the right thing.

Your LLM prompt and response are worthless.

Re: Google de-indexed Bear Blog and I don't know why

#67
post #61

Earlier quoted context omitted.

It's not even that. It's that "centralization is more efficient" is a big fat lie. If you look at the "centralized systems" they're... not actually technologically centralized, they're really just a monopolist that internally implements a distributed system. How do you think Google or Cloudflare actually work? One big server in San Francisco that runs the whole world, or lots of servers distributed all over?

I know exactly how they work, but they have a single entry point, as a customer you don't really care that the system is global, and they also have a single control plane, etc. Decisions are efficient if they need to be taken only once. The underlying architecture is irrelevant for the end user. Why do you think they're a monopoly in the first place? Obviously because they were more efficient than the competition and…

> Decisions are efficient if they need to be taken only once.

In other words, open source decentralized systems are the most efficient because you don't have to reduplicate a competitor's effort when you can just use the same code.

> Obviously because they were more efficient than the competition and network effects took care of the rest.

In most cases it's just the network effect, and whether it was a proprietary or open system in any given case is no more than the historical accident of which one happened to gain traction first.

> Having to make choices is a cost for the consumer

If you want an email address you can choose between a few huge providers and a thousand smaller ones, but that doesn't seem to prevent anyone from using it.

> until it switches to rent seeking

If it wasn't an open system from the beginning then that was always the end state and there is no point in waiting for someone to lock the door before trying to remove yourself from the cage.

Re: Google de-indexed Bear Blog and I don't know why

#68

Google search results have gone shit. I am facing some deindexing issues where Google is citing a content duplicate and picking a canonical URL itself, despite no similar content. Just the open is similar, but the intent is totally different, and so is the focus keyword. Not facing this issue in Bing and other search engines.

Yeah, Google search results are almost useless. How could they have neglected their core competence so badly?

Re: Google de-indexed Bear Blog and I don't know why

#70
Google search also favors large, well-known sites over newcomers. For sites that have a lot of competition, this is a real problem and leads to asymmetry and a chicken-and-egg problem. You are small/new, but you can't really be found, which means you can't grow enough to be found. At the same time, you are also disadvantaged because Google displays your already large competitors without any problems!
Post reply on HN