Live data from Hacker News

What happened to TheNumbers.com

stephenfollows.com

191–200 of 218 posts

Re: What happened to TheNumbers.com

#191

Earlier quoted context omitted.

> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem. > In other words: it sucks when people are using your work in a manner that you find offensive That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by…

This is absolutely fine, until some shitheel decides they're going to continually scrape your website, thousands of requests per second, and knock it offline constantly, so they don't have to cache even a single thing, or write better scrapers. They'll just suck up your bandwidth for it. After all, it's you that's paying, not them. Search engines were absolutely fine. They were respectful of the sources they scraped.…

Major AI companies aren't indiscriminately scrapping any more than search engines were (which themselves caused some problems in the past too, until that got sorted out). Your gripe is with SEO people and content marketers.

I'd be more sympathetic if it was really about this, though. However, the majority of the complaints seems to come from people personally offended by the possibility of their content, which they claimed was published for everyone to benefit, ending up as training data, and thus actually benefiting humanity, at scale far beyond the original publication ever would. Which is the Dog in the Manger attitude I point out.

Re: What happened to TheNumbers.com

#192
post #173

One thing I’m wondering is whether we need an open source set of technical patterns and libraries for dealing with this changing traffic mix. Millions of small sites and creators don’t have the ability to design their own protections against large scale automated access. If useful content now attracts aggressive crawler traffic, many sites will be too expensive or unreliable to run. Is part of the answer a community…

Yes, a community effort against the botnets would be great. The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem. It is the "bad bots" which pretend to be real users and hide behind residential proxies, and so are almost impossible to block at the moment, which are the problem. Given that there are big companies openly (i.e. on the clearweb, not even darkweb) selli…

> The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem.

And then they get blocked, which is a problem. In particular, the modern agentic AI tools interacting with web services to fulfill user queries - they are acting as user agents, and they should not be discriminated against.

So I'd say the first pattern that needs to be broadly adopted is non-discrimination of user agents.

But of course we've tried that in the past, the whole problem is that non-browser user agents == end-user automation, which is anathema to pretty much every on-line business out there, as money made online is primarily conditioned on users wasting their own lives on interacting with services directly.

Re: What happened to TheNumbers.com

#193

"One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products." I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way. I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anyth…

Why would you offer data for free if people are going to pay someone else to access it?

The whole internet is about to experience USA-level greed because people are too dumb not to use AI. Just like Google killed small sites because people were too lazy to look at page 2.

Re: What happened to TheNumbers.com

#194
post #158

Earlier quoted context omitted.

If it's the case that prediction market traders were trying to front-run the market outcomes, the prediction markets should put the fix on their end by locking down trading before a market closes with some time buffer to prevent this.

Why should that be necessary? You can unilaterally stop trading 'before a market closes with some time buffer to prevent this.' No need for centralised action.

Because markets succeed or fail based on perceived fairness.

It is in a prediction market’s best interest to not become the place where you go to get fleeced.

Today they’re the Wild West and run like the early days of darknet markets, but if the

Re: What happened to TheNumbers.com

#195

Earlier quoted context omitted.

>But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way. I dont get this. You provided something for free to help people, but dont want to do that anymore because it might go into training data and help many many more people? So far LLMs have been loan funded donations of loss leadin…

I think the sheer magnitude of the economics have made the scales fall from a lot of people's eyes. For decades people put stuff on the internet for free on the assumption it was "not worth" anything. It turns out that as soon as that commons can be enclosed, we can marshal hundreds of dollars for every single living human, to pay for this commons to be repackaged. The money is there, and we're happy to spend it, we…

So basically the “information is free, encyclopedias are expensive” phenomenon from the pre-internet days? Collection, collation, and distribution have more and different qualitative value than the sum of the individual bits.

Re: What happened to TheNumbers.com

#196

Earlier quoted context omitted.

This is absolutely fine, until some shitheel decides they're going to continually scrape your website, thousands of requests per second, and knock it offline constantly, so they don't have to cache even a single thing, or write better scrapers. They'll just suck up your bandwidth for it. After all, it's you that's paying, not them. Search engines were absolutely fine. They were respectful of the sources they scraped.…

Major AI companies aren't indiscriminately scrapping any more than search engines were (which themselves caused some problems in the past too, until that got sorted out). Your gripe is with SEO people and content marketers. I'd be more sympathetic if it was really about this, though. However, the majority of the complaints seems to come from people personally offended by the possibility of their content, which they c…

The problem is the rat race for the pot of gold at the end of the rainbow.

A similar thing happened when crypto was in ascendence. Web pages were getting crypto-miners injected into them. Celebrities were shilling NFTs. Everyone and their dog was on the make. The thought of riches broke the minds of millions.

The same is happening with AI. Whether it's the "major AI companies" or millions of self-interested also-rans with fewer moral scruples scraping, the problem still exists. The problem will continue to exist until you can't conceiveably make money by scraping like a bastard. If everybody identified themselves up front and respected robots.txt, there would not be a problem. But they don't, and they don't, and they pummel websites for no fucking reason, and they don't care, and they won't stop.

Websites get hammered by millions of unique IP addresses from residential ISPs which happen to belong to botnets, none of them identifying themselves as a bot user agent, all just pretending to be some slightly out-of-date version of Chrome.

And it's not like the "major AI companies" hands are clean. Who bought up all the RAM and compute power? Who's building datacenters in the third-world parts of the USA and powering 24/7 them with diesel generators, or sucking up most of the power generator, and of course using up all the drinking water to cool their machines?

Everything in humanity and nature is there for the AI bros to use and dispose of, provided they come out on top. This is the attitude that needs to be taken down.

Re: What happened to TheNumbers.com

#197
post #153

Earlier quoted context omitted.

Maybe because they restated what they said instead of addressing to the previous commenter’s point.

The comment I’m speaking of was from a different user, and it’s giving a strong smell that both users are bots unfortunately. Maybe I’m wrong, maybe it’s a coincidence those exact strings of characters were formed in isolation from each other in the same thread. I hope I’m wrong, but it’s a red flag

I'm not a bot, and the other commenter was restating the question as the solution provided did not seem to actually solve the problem, just redesign the entire project (See comment from u/taneq).

We didn't form "those exact strings of characters in isolation from each other in the same thread" they are directly referencing my comment in a reply thread to said comment.

Re: What happened to TheNumbers.com

#198

Earlier quoted context omitted.

> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem. > In other words: it sucks when people are using your work in a manner that you find offensive That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by…

Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet. In fact, open source embraces this (hence the 'open'). But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.

> Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet.

A more pragmatic approach is to acknowledge that concerning one's self with how something is used once it has been released is an emotional drain. It is a bit much to suggest that someone agreed to something, even if that agreement is implicit, just because they released it.

> But people are free to not publish things or post things online with a more restrictive license.

Licences are meaningless unless you have the ability to enforce them (e.g. to sue). That's why so many companies are willing to ignore the terms of open source licenses. It's also why the attempts of enforcement that we do hear about are usually backed by a third party, rather than being done by the software developer themselves. Simply put, the individual developer (or even small project) trying to make a contribution to the community would be better served by not publishing (instead of using a restrictive license) if they are concerned about how their work is used.

I don't even know if there is a good way to resolve the problem. Consider something like a DMCA Takedown notice. It removes the administrative and legal overhead to copyright infringement, yet it is also easy to abuse. For example: businesses have weaponized it by using it against individuals. Perhaps my cynicism is taking over here, but I suspect any easily accessible mechanism for enforcement would be similarly abused.

Re: What happened to TheNumbers.com

#199
post #113

Earlier quoted context omitted.

Why would you produce content and have no one read it or visits your site, but OpenAI and Anthropic make millions from it? At some point it becomes stupid to just give free money to these companies when they steal literally all your traffic and content. As per the article, Anthropic sends 1 page view for every 38,000 views they get.

Very few people showed up to read stuff on the old web, and yet: It existed.

The "old web" hasn't existed in 30 years now. I was a part of the old web and if people in the early 90s knew that someone was profiting off their work it would have killed it immediately.

Re: What happened to TheNumbers.com

#200
post #158

Earlier quoted context omitted.

Why should that be necessary? You can unilaterally stop trading 'before a market closes with some time buffer to prevent this.' No need for centralised action.

Because markets succeed or fail based on perceived fairness. It is in a prediction market’s best interest to not become the place where you go to get fleeced. Today they’re the Wild West and run like the early days of darknet markets, but if the

Could you please rephrase your last sentence? I appears to be incomplete.
Post reply on HN