Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

371–380 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#371
post #55

Earlier quoted context omitted.

It's a honeypot. He's telling people openai doesn't respect robots.txt and just scrapes whatever the hell it wants.

Except the first thing openai does is read robots.txt. However, robots.txt doesn't cover multiple domains, and every link that's being crawled is to a new domain, which requires a new read of a robots .txt on the new domain.

> Except the first thing openai does is read robots.txt.

What good is reading it if it doesn't respect it

Re: Anyone got a contact at OpenAI. They have a spider problem

#372

Earlier quoted context omitted.

> You didn't really do the joke right. It summed up many comments and attitudes I see here on LLM's in under 30 words. In fact it's so clever it's one of the few posts here I'm certain wasn't produced by an LLM.

The factual elements work as a summary. But the entire aspect of just deserts, getting what you asked for and being unhappy, hypocrisy, it isn't there. So by implying that kind of thing, it doesn't work. It's almost clever.

Lordy you are nerds.

The point is that free and open also equals free to be a predator. The joke is that people always think free and open just means what is good for them but truly free is often free for a predator to eat it all and make it closed.

While you guys lecture and circle jerk about the joke you miss the real point.

Like fuck dudes are you this dense? Your comments are convincing me LLMs are smarter than people.

Re: Anyone got a contact at OpenAI. They have a spider problem

#373

Earlier quoted context omitted.

I mean, you can use Phi now and it outclasses anything else in its size. This isn’t some “it could happen” situation.

Oh I’m sure it works wonderfully for now. My point is about the inevitable future when _those_ models start to struggle. The phi approach doesn’t seem like breaking the ouroboros, it just feels like inserting another model/snake into the loop.

“Struggle” at what? Struggle to have enough data to get smarter? Struggle to perform RAG and find legitimate sources?

I don’t think that we are going to get big improvements in LLMs without architecture improvements that need less data, and the current generation of models appears to be good enough at creating content from data/knowledge to train any future architectures we have with better synthetic datasets. Fortunately we have already seen examples of both of these “in the lab” and will probably see commercially sized models using some of the techniques in the coming months.

Re: Anyone got a contact at OpenAI. They have a spider problem

#374
post #124

Earlier quoted context omitted.

No, because there’s no legal weight behind robots.txt. The second someone weaponizes robots.txt all the scrapers will just start ignoring it.

That’s how you weaponize it. Set things up to give endless/randomized/poisoned data to anybody that ignores robots.txt.

You mean human users? That is and always will be the dominant group of clients that ignore robots.txt.

What you’re talking about is an arms race wherein bots try to mimic human users and sites try to ban the bots without also banning all their human users.

That’s not a fight you want to pick when one of the bot authors also owns the browser that 63% of your users use, and the dominant site analytics platform. They have terabytes of data to use to train a crawler to act like a human, and they can change Chrome to make normal users act like their crawler (or their crawler act more like a Chrome user).

Shit, if Google wanted, they could probably get their scrapes directly from Chrome and get rid of the scraper entirely. It wouldn’t be without consequence, but they could.

Re: Anyone got a contact at OpenAI. They have a spider problem

#376

Isn’t it funny that all the “worthless” content out there on the internet is actually changing the world. Like how 4chan was mocked as being a cesspit for losers, but now everyone knows memes like Pepe the frog and Wojack from there. And like now this very comment and the billions of other comments on here, Reddit, Twitter, etc that are regarded as a “waste of time” are being used to train multi billion dollar compan…

Sharing it is often the only reason why it has value in the first place. If the authors of Gangnam Style or The Fox or any other silly viral thing hadn’t shared them they would have had zero value.

Re: Anyone got a contact at OpenAI. They have a spider problem

#377
post #124

Earlier quoted context omitted.

That’s how you weaponize it. Set things up to give endless/randomized/poisoned data to anybody that ignores robots.txt.

You mean human users? That is and always will be the dominant group of clients that ignore robots.txt. What you’re talking about is an arms race wherein bots try to mimic human users and sites try to ban the bots without also banning all their human users. That’s not a fight you want to pick when one of the bot authors also owns the browser that 63% of your users use, and the dominant site analytics platform. They ha…

It’s fairly trivial to treat Google’s crawler differently if you want. https://developers.google.com/search/docs/crawling-indexing/...

The point here is to poison the well for freeloaders like OpenAI not to actually prevent web crawlers. OpenAI will actually pay for access to good training data, don’t hand it over for free.

People don’t mindlessly click on things like terms of service crawlers are quite dumb. Little need for an arms race, as the people running these crawlers rarely put much effort into any one source.

Re: Anyone got a contact at OpenAI. They have a spider problem

#378

Earlier quoted context omitted.

I'm referring to his two works, the "Tractatus Logico-Philosophicus" and "Philosophical Investigations". There's a lot explored here, but Wittgenstein basically makes the argument that the natural logic of language—how we deduce meaning from terms in a context and naturally disambiguate the semantics of ambiguous phrases—is different from the sort of formal propositional logic that forms the basis of western philosop…

Wait, are you saying this something you read in both the Tractatus and the PI? They are quite opposed as texts! That's kinda why he wrote the PI at all.. I don't think Wittgenstein would agree, first of all, that there is a "natural logic" to language. At least in the PI, that kind of entity--"the natural logic of language"--is precisely the kind of weird and imprecise use of language he is trying to expose. Even mor…

> They are quite opposed as texts!

The texts are quite different, this is true, but I don't find them contradictory. Whereas Tractatus was almost a facetious or flippant rejection of the millenia-long project to agree on a philosophical subset of language suitable for rigorous philosophy (although it continues today in the form of analytical philosophy), PI basically says "well we don't need to throw the baby out with the bath water", which I think is a fantastically mature response to a flawed tool that's still the best we have to reason about the universe. So: not contradictory in evaluation of fundamental compatibility of non-formal language for the formal needs of propositional philosophy, but perhaps contradictory in implied reaction to this realization.

Re: Anyone got a contact at OpenAI. They have a spider problem

#379
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

I prefer "error" but would be OK with "mistake".

The problem with "hallucination" and "confabulation" is that they both imply a consciousness.

Re: Anyone got a contact at OpenAI. They have a spider problem

#380
post #377

Earlier quoted context omitted.

You mean human users? That is and always will be the dominant group of clients that ignore robots.txt. What you’re talking about is an arms race wherein bots try to mimic human users and sites try to ban the bots without also banning all their human users. That’s not a fight you want to pick when one of the bot authors also owns the browser that 63% of your users use, and the dominant site analytics platform. They ha…

It’s fairly trivial to treat Google’s crawler differently if you want. https://developers.google.com/search/docs/crawling-indexing/... The point here is to poison the well for freeloaders like OpenAI not to actually prevent web crawlers. OpenAI will actually pay for access to good training data, don’t hand it over for free. People don’t mindlessly click on things like terms of service crawlers are quite dumb. Little…

It’s trivial to treat it differently, but doing so runs the risk of being accused of cloaking and getting banned from Google’s index: https://developers.google.com/search/docs/essentials/spam-po...

> The point here is to poison the well for freeloaders like OpenAI not to actually prevent web crawlers. OpenAI will actually pay for access to good training data, don’t hand it over for free.

Sure, and they’ll pay the scrapers you haven’t banned for your content, because it costs those scrapers $0 to get a copy of your stuff so they can sell it for far less than you.

> People don’t mindlessly click on things like terms of service crawlers are quite dumb. Little need for an arms race, as the people running these crawlers rarely put much effort into any one source.

The bots are currently dumb _because_ we don’t try to stop them. There’s no need for smarter scrapers.

Watch how quickly that changes if people start blocking bots enough that scraped content has millions of dollars of value.

At the scale of a company, it would be trivial to buy request log dumps from one of the adtech vendors and replay them so you are legitimately mimicking a real user.

Even if you are catching them, you also have to be doing it fast enough that they’re not getting data. If you catch them on the 1,000th request, they’re getting enough data that it’s worthwhile for them to just rotate AWS IPs when you catch them.

Worst case, they just offer to pay users directly. “Install this addon. It will give you a list of URLs you can click to send their contents to us. We’ll pay you $5 for every thousand you click on.” There’s a virtually unlimited supply of college students willing to do dumb tasks for beer money.

You can’t price segment a product that you give away to one segment. The segment you’re trying to upcharge will just get it for cheap from someone you gave it to for free. You will always be the most expensive supplier of your own content, because everyone else has a marginal cost of $0.

Post reply on HN