Live data from Hacker News

What happened to TheNumbers.com

stephenfollows.com

171–180 of 218 posts

Re: What happened to TheNumbers.com

#171
post #8

I think people are missing one of the points of the article. It's not just that agents are hammering the site, it's that there might be lurking vulnerabilities that allow malicious usage, which is why it went down, then came back up with a fraction of the data and a reduced design. The article says (speculates?) that malicious users are trying to get privileged access for an edge in prediction market betting. From th…

If it's the case that prediction market traders were trying to front-run the market outcomes, the prediction markets should put the fix on their end by locking down trading before a market closes with some time buffer to prevent this.

That's assuming the prediction markets all play fair and by the same rules. If Polymarket were to do that, its competition could choose not to. Since these exchanges are not tied to any one country (or jurisdiction), its users would jump ship. Because the users don't care.

See also stock market trading, where companies would go to great lengths to shave nanoseconds off of getting information and putting in trades. Now apply this min/maxing to worldwide, unregulated and anonymous.

This is what tech libertarians / cryptobros want.

Re: What happened to TheNumbers.com

#172
post #45

Earlier quoted context omitted.

> But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way. The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem. > And I can see why maintainers of sites like this, or other free but i…

> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem. > In other words: it sucks when people are using your work in a manner that you find offensive That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by…

Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet. In fact, open source embraces this (hence the 'open').

But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.

Re: What happened to TheNumbers.com

#173

One thing I’m wondering is whether we need an open source set of technical patterns and libraries for dealing with this changing traffic mix. Millions of small sites and creators don’t have the ability to design their own protections against large scale automated access. If useful content now attracts aggressive crawler traffic, many sites will be too expensive or unreliable to run. Is part of the answer a community…

Yes, a community effort against the botnets would be great.

The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem. It is the "bad bots" which pretend to be real users and hide behind residential proxies, and so are almost impossible to block at the moment, which are the problem.

Given that there are big companies openly (i.e. on the clearweb, not even darkweb) selling access to these residential proxy botnets of compromised smart TVs[0] and mobile phones and other devices, can we not simply get a database of the these IPs and block access from them? I would venture that almost every single one of those residential users are unaware that they have devices in their homes which have been compromised and are being abused in this way, so if they were to start seeing messages from more and more sites along the lines of "Access to this site has been blocked because unusual traffic has been detected from your computer network. Please check all devices on your network and remove any malware which may be routing this traffic." then maybe we could start addressing the problem at the source.

[0] https://news.ycombinator.com/item?id=49000864

Re: What happened to TheNumbers.com

#174
post #154

AI companies really do socialize the costs and privatize the profits. Sites like The Numbers have to take on the cost of surviving the AI onslaught and the AI companies return nothing back to them.

To a point; as the post mentions, they added instructions for LLMs to their website and increased their licensing inquiries tenfold (the article didn't state anything about actual licensing payments being made though).

I don't think there would be any issues per se if the scrapers just paid licensing fees to get the good / complete data. But the issue was that they started to try and find exploits to get to data earlier.

Re: What happened to TheNumbers.com

#175

Earlier quoted context omitted.

> Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way. I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.

My understanding of the GPL is that it currently prohibits the training of LLMs, no need to explicitly mention them. MIT does allow it, however.

Exactly. Just like using LibGen is prohibited.

Re: What happened to TheNumbers.com

#176

Just FYI, the bigger companies all allow you to block crawlers via robots.txt: # Block Anthropic (Claude) User-agent: ClaudeBot Disallow: / User-agent: Claude-SearchBot Disallow: / User-agent: Claude-User Disallow: / # Block OpenAI (ChatGPT) User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Disallow: / # Block Perplexity User-agent: PerplexityBot Disallow: / # Block Google's AI Training User-agent: Google-Exte…

This only works for crawlers that play fair, unfortunately there's many that don't. Including ones that use consumer devices like TVs [0] and mobile phone apps that offer incentives to consumers (if they're even open about it) to use their internet connection to e.g. crawl websites. There's probably an army of hacked toothbrushes and the like doing the same thing.

[0] https://news.ycombinator.com/item?id=49000864

Re: What happened to TheNumbers.com

#177
Only solution I see is to ban user that abuse the system. Too much requests, too much bandwidth and you get into the blacklist and are cutoff from the human side of the internet.

With the amount of money AI labs are burning, someone should just set up some infra to host the data they want and charge for access to it, instead of abusing the goodwill of legitimate pages that offer it for free.

Re: What happened to TheNumbers.com

#178
post #45

Earlier quoted context omitted.

> But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way. The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem. > And I can see why maintainers of sites like this, or other free but i…

> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem. > In other words: it sucks when people are using your work in a manner that you find offensive That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by…

This is absolutely fine, until some shitheel decides they're going to continually scrape your website, thousands of requests per second, and knock it offline constantly, so they don't have to cache even a single thing, or write better scrapers. They'll just suck up your bandwidth for it. After all, it's you that's paying, not them.

Search engines were absolutely fine. They were respectful of the sources they scraped. There's now a bunch of money-obsessed shitheels doing whatever they can, as blatantly and as lazily as they can, because they believe they'll become rich. Every one of these fuckers needs to be dead or in jail before the world becomes whole again.

If nobody takes action against these fuckers, everything you care about will just nope out of existence, and these fuckers will be all you have left.

Re: What happened to TheNumbers.com

#179

Earlier quoted context omitted.

A change in quantity can become a change in quality.

I believe that it was Stalin who phrased it most eloquently: Quantity has a quality all its own.

Not Stalin: https://klangable.com/blog/quantity-has-a-quality-all-its-ow...

Re: What happened to TheNumbers.com

#180

"One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products." I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way. I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anyth…

> Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way. I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.

If LLM training is found to be a fair use (looks likely), no license will help.
Post reply on HN