Live data from Hacker News

US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

storage.courtlistener.com

321–330 of 1001 posts

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#321
post #280

In terms of fairness, competition and monopolies is there a chart that shows how much tax payer funding each search engine has received upon creation, annually and indirectly ? e.g. donating NASA hangers for server hosting and experiments, heavily discounted real estate and land, tax breaks for power, etc... Put another way, who has the biggest monopoly on direct and indirect tax-payer funding?

> Put another way, who has the biggest monopoly on direct and indirect tax-payer funding?

The Pentagon.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#322
The really uncompetitive behaviour started when Google removed the search string from search links. That killed 3rd party (and home-grown) analytics, which in turn facilitated large scale tracking of users from site with analytics to site with analytics.

If you wanted to know how your keywords were performing you had to use Google Analytics.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#323
post #23

> As a result, YC has an interest in ensuring that U.S. technology markets are free from anticompetitive barriers to entry and expansion. This part reads like a suggestion to loosen anti-competitive/antitrust law.

Why? The point of antitrust is to promote market fairness.

I don't trust YC very much, but I do trust they want a share of the pie. And they're not wrong that Google has monopolized and stagnated search. I think you're reading too much into that sentence?

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#324

Earlier quoted context omitted.

Crawling the internet is a natural monopoly. Nobody wants an endless stream of bots crawling their site, so googlebot wins because they’re the dominant search engine. It makes sense to break that out so everyone has access to the same dataset at FRAND pricing. My heart just wants Google to burn to the ground, but my brain says this is the more reasonable approach.

Google search is a monopoly not because of crawling. It's because of the all the data it knows about website stats and user behavior. Original Google idea of ranking based on links doesn't work because it's too easily gamed. You have to know what websites are good based on user preferences and that's where you need to have data. It's impossible to build anything similar to Google without access to large amounts of us…

Page ranking sounds like a perfect application of artificial intelligence.

If China can apply it for total information awareness on their population, Google can apply it on page reliability

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#325

Earlier quoted context omitted.

Which is why, like the 'monopoly on violence' the government should also be funding a _lot more research_. It should be at, or partnered with, higher learning institutions and since it's public funded all of the results should be free to use*. I'm willing to entertain the idea of: Free use for people and corporations within the country/countries that funded research, everyone else pays compulsory license fees.

But public funded research isn’t “free to use.” In many cases, you can’t even read it without paying a scientific journal for a subscription. See the Bayh-Dole Act as well: universities can patent discoveries from federally funded research.

Publications with public funding have already escaped the paywall, partially as of 2013 and completely as of this year:

https://par.nsf.gov/

https://pmc.ncbi.nlm.nih.gov/

https://ospo.gwu.edu/overview-us-policy-open-access-and-open...

https://www.nih.gov/about-nih/who-we-are/nih-director/statem...

https://www.coalition-s.org/plan_s_principles/

The intent of the Bayh-Dole Act was to deal with a perceived problem of government-owned patents being investor-unfriendly. At the time the government would only grant non-exclusive licenses, and investors generally want exclusivity. That may have been the actual problem, moreso than who owned the patent. On the other hand, giving the actual inventors an incentive to commercialize their work should increase their productivity and the chance that the inventions actually get used.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#326

Earlier quoted context omitted.

> Crawling the internet is a natural monopoly. How so? A caching proxy costs you almost nothing and will serve thousands of requests per second on ancient hardware. Actually there's never been a better time in the history of the Internet to have competing search engines since there's never been so much abundance of performance, bandwidth, and software available at historic low prices or for free.

Costs almost nothing, but returns even less.* There are so many other bots/scrapers out there that literally return zero that I don’t blame site owners for blocking all bots except googlebot. Would it be nice if they also allowed altruist-bot or common-crawler-bot? Maybe, but that’s their call and a lot of them have made it on a rational basis. * - or is perceived to return

> that I don’t blame site owners for blocking all bots except googlebot.

I doubt this is happening outside of a few small hobbyist websites where crawler traffic looks significant relative to human traffic. Even among those, it’s so common to move to static hosting with essentially zero cost and/or sign up for free tiers of CDNs that it’s just not worth it outside of edge cases like trying to host public-facing Gitlab instances with large projects.

Even then, the ROI on setting up proper caching and rate limiting far outweighs the ROI on trying to play whack-a-mole with non-Google bots.

Even if someone did go to all the lengths to try to block the majority of bots, I have a really hard time believing they wouldn’t take the extra 10 minutes to look up the other major crawlers and put those on the allow list, too.

This whole argument about sites going to great lengths to block search indexers but then stopping just short of allowing a couple more of the well-known ones feels like mental gymnastics for a situation that doesn’t occur.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#327

Meanwhile, YC has happily and excitedly fed it's start-ups to Google over the years. So pretty much "We don't want google to develop new things, we want them to have buy those from us"

> Meanwhile, YC has happily and excitedly fed it's start-ups to Google over the years.

I'm curious what you've seen or heard that led you to that conclusion? It's the opposite of correct.

YC supports what founders want, including if they want to sell to $BigCo, but such outcomes are hardly successes for YC. YC's success depends on outlier companies growing much larger than that.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#328

It's good for YC to do this and will benefit every startup in the long run. Google has been one of the sources of the AI boom, and provides liquidity by acquiring startups. But as YC argues they've monopolised distribution channels to the point where you need to go through the Google toll booth every time you want to access the market. This tax on founders to reach their audience makes many types of businesses unsust…

Of course it is good for YC to do this. They have significant investments in OpenAI both directly and indirectly through the countless startups they've funded whose core is OpenAI.

And it's ridiculous to act like (a) you are forced to go through Google to access the 'market' and (b) that this is somehow unusual or untoward. They are an advertising company and not the only one.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#329
Make no mistake. This is, first and foremost, a big, for-profit corporation fighting a bigger, for-profit corporation, for its own financial interests. Nevertheless, we may stand to benefit, if only incidentally.

In particular, if the legal authorities start to unwind Google, I actually think Chrome and Android are more important to wall off or spin out than anything advertising or AI related.

Re: US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]

#330
post #262

Earlier quoted context omitted.

Are sites really that averse to having a few more crawlers than they already do? It would seem that it’s only a monopoly insofar as it’s really expensive to do and almost nobody else thinks they can recoup the cost.

A few? We routinely are fighting off hundreds of bots at any moment. Thousands and Thousands per day, easily. US, China, Brazil from hundreds of different IPs, dozens of different (and falsified!) user agents all ignoring robots.txt and pushing over services that are needed by human beings trying to get work done. EDIT: Just checked our anubis stats for the last 24h CHALLENGE: 829,586 DENY: 621,462 ALLOW: 96,810 This…

This seems like two different issues.

One is, suppose there are a thousand search engine bots. Then what you want is some standard facility to say "please give me a list of every resources on this site that has changed since " so they can each get a diff from the last time they crawled your site. Uploading each resource on the site to each of a thousand bots once is going to be irrelevant to a site serving millions of users (because it's a trivial percentage) and to a site with a small amount of content (because it's a small absolute number), which together constitute the vast majority of all sites.

The other is, there are aggressive bots that will try to scrape your entire site five times a day even if nothing has changed and ignore robots.txt. But then you set traps like disallowing something in robots.txt and then ban anything that tries to access it, which doesn't affect legitimate search engine crawlers because they respect robots.txt.

Post reply on HN