Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

381–390 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#381
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

Reminds me of a joke Three logicians walk into a bar. The bartender says "what'll it be, three beers?" The first logician says "I don't know". The second logician says "I don't know". The third logician says "Yes".

I do love a good joke, but this one falls a bit flat.

Logically speaking, the second bar tender could have thought to himself "no I don't want any beer, but one of these two other guys may want to double fist" and so there is really no way for the third logician to answer in the affirmative.

Re: Anyone got a contact at OpenAI. They have a spider problem

#382
post #377

Earlier quoted context omitted.

It’s fairly trivial to treat Google’s crawler differently if you want. https://developers.google.com/search/docs/crawling-indexing/... The point here is to poison the well for freeloaders like OpenAI not to actually prevent web crawlers. OpenAI will actually pay for access to good training data, don’t hand it over for free. People don’t mindlessly click on things like terms of service crawlers are quite dumb. Little…

It’s trivial to treat it differently, but doing so runs the risk of being accused of cloaking and getting banned from Google’s index: https://developers.google.com/search/docs/essentials/spam-po... > The point here is to poison the well for freeloaders like OpenAI not to actually prevent web crawlers. OpenAI will actually pay for access to good training data, don’t hand it over for free. Sure, and they’ll pay the scr…

Google doesn’t care what you do to other crawlers that ignore your TOS. This isn’t a theoretical situation it’s already going on. Crawling is easy enough to “block” there’s court cases on this stuff because this is very much the case where the defense wins once they devote fairly trivial resources to the effort.

And again blocking should never be the goal poisoning the well is. Training AI on poisoned data is both harder to detect and vastly more harmful. A price compared tool is only as good as the actual prices it can compare etc.

Re: Anyone got a contact at OpenAI. They have a spider problem

#383

Earlier quoted context omitted.

It's a stretch to expect a human initiated action to abide by robot.txt. Also, once you click on a link in chrome it's pretty much all robot parsed and rendered from there as well..

At bottom, all robot actions are human initiated.

Exactly, so they should ignore robots.txt !

Re: Anyone got a contact at OpenAI. They have a spider problem

#384

Earlier quoted context omitted.

The factual elements work as a summary. But the entire aspect of just deserts, getting what you asked for and being unhappy, hypocrisy, it isn't there. So by implying that kind of thing, it doesn't work. It's almost clever.

Lordy you are nerds. The point is that free and open also equals free to be a predator. The joke is that people always think free and open just means what is good for them but truly free is often free for a predator to eat it all and make it closed. While you guys lecture and circle jerk about the joke you miss the real point. Like fuck dudes are you this dense? Your comments are convincing me LLMs are smarter than p…

> truly free is often free for a predator to eat it all and make it closed

That doesn't apply when the word "open" is helping clarify.

A predator making open things closed isn't people asking for the wrong thing and getting egg on their face, it's people not getting what they asked for at all through no fault of their own.

Especially because this same kind of crawling works on a non-open internet.

> circle jerk

Do you say that any time someone disagrees with you, or is there something here in specific that you're reacting to?

Re: Anyone got a contact at OpenAI. They have a spider problem

#385
post #106

Earlier quoted context omitted.

Accessing a directly referenced page is common in order to receive the noindex header and/or meta tag, whose semantics are not implied by “Disallow: /” And then all the links are to external domains, which aren't subject to the first site's robots.txt

This is a moderately persuasive argument. Although the crawler should probably ignore all the html body. But it does feel like a grey area if I accept your first pint.

You've been able to convince me to accept his second pint. Friday it is.

Re: Anyone got a contact at OpenAI. They have a spider problem

#386

Earlier quoted context omitted.

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

I prefer "error" but would be OK with "mistake". The problem with "hallucination" and "confabulation" is that they both imply a consciousness.

I'm not sure that's correct. Hallucination definitely implies some sort of connection to reality, confabulation does not, it implies some kind of hard to detect error in stitching memories together coherently

Re: Anyone got a contact at OpenAI. They have a spider problem

#387
post #67
post #55

Earlier quoted context omitted.

Except the first thing openai does is read robots.txt. However, robots.txt doesn't cover multiple domains, and every link that's being crawled is to a new domain, which requires a new read of a robots .txt on the new domain.

> Except the fiist thing openai does is read robots.txt. Then they should see the "Disallow: /" line, which means they shouldn't crawl any links on the page (because even the homepage is disallowed). Which means they wouldn't follow any of the links to other subdomains.

This robots.txt has Disallow rule commented out:

    # buzz off
    #User-agent: GPTBot
    #Disallow: /

Re: Anyone got a contact at OpenAI. They have a spider problem

#388
post #343

Earlier quoted context omitted.

why is it three logicians? wouldn't it work with just two?

I recently heard this explained ( ) in the following way: three is the smallest number where you can set up an expectation (with the first two) and then break it. This is why three is such a common number, not just in jokes but in all sorts of story-telling. ( ) In a lecture by the mathematician & author Sarah Hart.

This guy footnotes.

Re: Anyone got a contact at OpenAI. They have a spider problem

#389

Earlier quoted context omitted.

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

I prefer "error" but would be OK with "mistake". The problem with "hallucination" and "confabulation" is that they both imply a consciousness.

Technically, then everything it's doing are errors just with varying degrees of precision.

It's still just probabilistic babble. We were doing this with markov chains.

Re: Anyone got a contact at OpenAI. They have a spider problem

#390

Earlier quoted context omitted.

Lordy you are nerds. The point is that free and open also equals free to be a predator. The joke is that people always think free and open just means what is good for them but truly free is often free for a predator to eat it all and make it closed. While you guys lecture and circle jerk about the joke you miss the real point. Like fuck dudes are you this dense? Your comments are convincing me LLMs are smarter than p…

> truly free is often free for a predator to eat it all and make it closed That doesn't apply when the word "open" is helping clarify. A predator making open things closed isn't people asking for the wrong thing and getting egg on their face, it's people not getting what they asked for at all through no fault of their own. Especially because this same kind of crawling works on a non-open internet. > circle jerk Do yo…

No just to the misunderstanding of what open is. So many people thing free and open also has rules that protect them.

It does not.

Post reply on HN