Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

31–40 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#32
post #22
post #14

Earlier quoted context omitted.

Arachnid here… What am I looking at? The intent is to waste the resources of crawlers by just making the web larger?

What am I looking at? I'd say go ahead and inject it with digestive enzymes and then report findings.

No no. First tightly wrap it in silk. Then digestive enzymes. Good hygiene, eh?

Re: Anyone got a contact at OpenAI. They have a spider problem

#36
post #15

I’d let them do their thing, why not?! They want the internet? This is the real internet. It looks like he doesn’t really care that much that they’re retrieving millions of pages, so let them do their thing…

> It looks like he doesn’t really care that much that they’re retrieving millions of pages

It impacts the performance for the other legitimate users of that web farm ;)

Re: Anyone got a contact at OpenAI. They have a spider problem

#38

I'm more interested in what that content farm is for. It looks pointless, but I suspect there's a bizarre economic incentive. There are affiliate links, but how much could that possibly bring in?

It'd say it's more like a honeypot for bots. So pretty similar objectives.

Re: Anyone got a contact at OpenAI. They have a spider problem

#39
post #10

Earlier quoted context omitted.

I think he's saying that it's not a problem for him, but for OpenAI?

Yup, that's my impression as well. He's just nice to let OpenAI they have a problem. Usually this should be rewarded with a nice "hey, u guys have a bug" bounty because not long time ago some VP from OpenAI was lamenting that training their AI is, and it's his direct quote, "eye watering" cost (the order was millions of $$ per second).

I would be a little sceptical about that figure. 3 million dollars per second is around the world GDP.

I get it, AI training is expensive, but I don't believe it's that expensive

Re: Anyone got a contact at OpenAI. They have a spider problem

#40

Isn’t the legality of web scraping still..disputed? There’s been a few projects I’ve wanted to work on involving scraping, but the idea that the entire thing could be shut down with legal threats seems to make some of the ideas infeasible. It’s strange that OpenAI has created a ~$80B company (or whatever it is) using data gathered via scraping and as far as I’m aware there haven’t been any legal threats. Was there so…

> It’s strange that OpenAI has created a ~$80B company (or whatever it is) using data gathered via scraping

Like Google and many others.

Post reply on HN