Live data from Hacker News

US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

storage.courtlistener.com

351–360 of 1001 posts

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#351

Earlier quoted context omitted.

Google search is a monopoly not because of crawling. It's because of the all the data it knows about website stats and user behavior. Original Google idea of ranking based on links doesn't work because it's too easily gamed. You have to know what websites are good based on user preferences and that's where you need to have data. It's impossible to build anything similar to Google without access to large amounts of us…

Page ranking sounds like a perfect application of artificial intelligence. If China can apply it for total information awareness on their population, Google can apply it on page reliability

I'm fairly certain many people have already tried to apply magical AI pixie dust to this problem. Presumably it isn't so simple in practice.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#353
post #346

TL;DR: YC just filed an amicus brief in the Google-search antitrust case. They tell the judge that (1) Google’s default-search/pay-to-play deals crushed the “kill-zone” around search, keeping VCs away; and (2) the coming wave of AI-search/agentic tools will suffer the same fate unless the court imposes forward-looking remedies—open Google’s index, ban exclusive data+distribution deals, bar self-preferencing, add anti…

Summarized by o3:

1. Why YC cares

- VC “kill-zone.” YC says Google’s decade of default-search contracts (Apple, carriers, OEMs) froze half of all U.S. search queries, scaring investors away from search/AI startups.

- AI inflection point. Generative/query-based/agentic AI could disrupt search—but only if newcomers can reach users and train on data Google hoards. Without action, Google will “pull the ladder up” again.

2. What YC wants the court to order

- Open index & dataset access. Force Google to license its search index + anonymized click/embedding data on fair, reasonable terms so rivals can build ranking stacks + AI models.

- No self-preferencing in AI results. Google can’t boost Gemini-style tools or demote rivals. No exclusive AI-training corpora access either.

- Ban pay-to-play defaults. Outlaw “billions to be the default” (search, voice, browser, OS, car). No payments for choice-screen placement.

- Anti-circumvention & retaliation guardrails. Independent monitoring, fast dispute resolution, steep fines, and—if Google cheats—possible Android spinoff.

3. Historical playbook YC cites

- 1956 AT&T consent decree opened Bell Labs patents → semiconductor boom.

- 2001 Microsoft browser decree → Firefox, Chrome, Google itself.

- 2023–24 Nvidia-Arm block → both companies exploded in AI.

- YC says same pattern can unlock “the next Stripe/Airbnb—but for search/AI.”

4. Why HN should care

- Open Index ≈ ultimate dev API. Lets you build retrieval-augmented AI agents without a nation-state’s crawl budget.

- Distribution shake-up. Killing default deals revives mobile/browser competition; could birth real alt-search on phones.

- VC signal. YC telling a judge “give us a level field and we’ll bankroll challengers” means real capital is ready.

- Policy trend. Regulators now want to pre-wire markets (index access, AI data parity, Android contingency) before the next moat forms.

5. Bottom line: This isn’t about a fine. It’s about cracking open the data + distribution bottlenecks that froze search since 2009. If Judge Mehta adopts YC/DOJ’s plan, the door opens for real search/AI innovation—and VCs are ready to sprint through it.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#354

Earlier quoted context omitted.

> Crawling the internet is a natural monopoly. How so? A caching proxy costs you almost nothing and will serve thousands of requests per second on ancient hardware. Actually there's never been a better time in the history of the Internet to have competing search engines since there's never been so much abundance of performance, bandwidth, and software available at historic low prices or for free.

Costs almost nothing, but returns even less.* There are so many other bots/scrapers out there that literally return zero that I don’t blame site owners for blocking all bots except googlebot. Would it be nice if they also allowed altruist-bot or common-crawler-bot? Maybe, but that’s their call and a lot of them have made it on a rational basis. * - or is perceived to return

> that I don’t blame site owners for blocking all bots except googlebot

I run a number of sites with decent traffic and the amount of spam/scam requests outnumbers crawling bots 1000 to 1.

I would guess that the number of sites allowing just Googlebot is 0.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#355

Earlier quoted context omitted.

Costs almost nothing, but returns even less.* There are so many other bots/scrapers out there that literally return zero that I don’t blame site owners for blocking all bots except googlebot. Would it be nice if they also allowed altruist-bot or common-crawler-bot? Maybe, but that’s their call and a lot of them have made it on a rational basis. * - or is perceived to return

> that I don’t blame site owners for blocking all bots except googlebot. I doubt this is happening outside of a few small hobbyist websites where crawler traffic looks significant relative to human traffic. Even among those, it’s so common to move to static hosting with essentially zero cost and/or sign up for free tiers of CDNs that it’s just not worth it outside of edge cases like trying to host public-facing Gitla…

> sites going to great lengths to block search indexers

That's not it. They're going to great lengths to block all bot traffic because of abusive and generally incompetent actors chewing through their resources. I'll cite that anubis has made the front page of HN several times within the past couple months. It is far from the first or only solution in that space, merely one of many alternatives to the solutions provided by centralized services such as cloudflare.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#356

Earlier quoted context omitted.

> Crawling the internet is a natural monopoly. How so? A caching proxy costs you almost nothing and will serve thousands of requests per second on ancient hardware. Actually there's never been a better time in the history of the Internet to have competing search engines since there's never been so much abundance of performance, bandwidth, and software available at historic low prices or for free.

You don't get to tell site owners what to do. The actual facts on the ground are that they're trying to block your bot. It would be nice if they didn't block your bot, but the other, completely unnatural and advertising-driven, monopoly of hosting providers with insane per-request costs makes that impossible until they switch away.

> The actual facts on the ground are that they're trying to block your bot

Based on what evidence.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#357

The solution proposed by Kagi—separate the search index from the rest of Google—seems to make the most sense. Kagi explains it more here: https://blog.kagi.com/dawn-new-era-search

At Blekko we advocated for this as well. Google has two interlocked monopolies, one is the search index and the other is their advertising service. We often joked that if Google reasonable and non-discriminatory priced access to their index, both to themselves and to others, AND they allowed someone to put what ever ads they wanted on those results. That change the landscape dramatically. Google would carve out their…

You mean like a white label search engine? Customized with settings?

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#358

Earlier quoted context omitted.

And Netflix shouldn’t have been able to extend its monopoly in shipping DVDs to streaming

That's not comparable. Netflix did not have a monopoly on "shipping DVDs". Plenty of retailers, online or brick-and-mortar, did that at the same time Netflix did, and video rental places were still going. Some of which would deliver movies (the local Marcos near me had a deal with a local Family Video where if you bought a large pizza they would bring you both the pizza and a movie of your choice from Family Video, a…

And there are other traditional search engines - including one run by another 1 trillion dollar+ market cap company.

That’s not to mention ChatGPT and other LLMs have built in search.

Should Google not evolve with the times?

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#359

Earlier quoted context omitted.

Assuming the simplified diagram of Google’s architecture, sure, it looks like you’re just splitting off a well-isolated part, but it would be a significant hardship to do it in reality. Why not also require Apple to split off only the phone and messaging part of its iPhone, Meta to split off only the user feed data, and for the U.S. federal government to run only out of Washington D.C.? This isn’t the breakup of AT&T…

Google killed Google. They should not have decided to become evil. Search can easily be removed, G Suite should be separate too.

> Search can easily be removed

This strikes me like "two easy steps to draw an owl. First draw the head, then draw the body". I generally support some sort of breakup, but hand waving the complexities away is not going to do anybody any good

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#360

The solution proposed by Kagi—separate the search index from the rest of Google—seems to make the most sense. Kagi explains it more here: https://blog.kagi.com/dawn-new-era-search

That's like asking the foxes how the farmer should manage his chickens. Kagi is a (wannabe) competitor. Likewise, YC's interest here is in making money by having viable startups and having them acquired.

I also don't think crawling the Web is the hard part. It's extraordinarily easy to do it badly [1] but what's the solution here? To have a bunch of wannabe search engines crawl Google's index instead?

I've thought about this and I wonder if trying to replicate a general purpose search engine here is the right approach or not. Might it not be easier to target a particular vertical, at least to start with? I refuse to believe Google cannot be bested in every single vertical or that the scale of the job can't be segmented to some degree.

[1]: https://stackoverflow.blog/2009/06/16/the-perfect-web-spider...

Post reply on HN