Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

161–170 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#162
post #141

Earlier quoted context omitted.

With Wittgenstein I think we see that "hallucinations" are a part of language in general, albeit one I could see being particularly vexing if you're trying to build a perfectly controllable chatbot.

This sounds interesting, could you give more detail on what you're referring to?

I would assume GP is talking about the fallibility of human memory, or perhaps about the meanings of words/phrases/aphorisms that drift with time. C.S. Lewis talks about the meaning of the word "gentleman" in one of his books; at first the word just meant "land owner" and that was it. Then it gained social significance and began to be associated with certain kinds of behavior. And now, in the modern register, its meaning is so dilute that it can be anything from "my grandson was well behaved today" or "what an asshole" depending on its use context.

Dunno. GP?

Re: Anyone got a contact at OpenAI. They have a spider problem

#163
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

With Wittgenstein I think we see that "hallucinations" are a part of language in general, albeit one I could see being particularly vexing if you're trying to build a perfectly controllable chatbot.

I don't remember Wittgenstein saying anything about that.

Re: Anyone got a contact at OpenAI. They have a spider problem

#164

Earlier quoted context omitted.

> I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs I mean, it's inherent to LLMs to be unable to answer "I don't know" as a result of not knowing the answer . An LLM never "doesn't know" the answer. But they'll gladly answer "I don't know" if that's statistically the most likely response, right? (Although current public offerings are probably trained against ev…

LLMs work at all because of the high correlation between the statistically most likely response and the most reasonable answer.

That's an explanation of why their answers can be useful, but doesn't relate to their ability to "not know" an answer

Re: Anyone got a contact at OpenAI. They have a spider problem

#165
post #136

Earlier quoted context omitted.

Your question may be a better fit for a different StackExchange site.

We prefer questions that can be answered, not merely discussed.

Feel like we're going to get dang on our case soon...

Re: Anyone got a contact at OpenAI. They have a spider problem

#166
post #141

Earlier quoted context omitted.

With Wittgenstein I think we see that "hallucinations" are a part of language in general, albeit one I could see being particularly vexing if you're trying to build a perfectly controllable chatbot.

This sounds interesting, could you give more detail on what you're referring to?

I'm referring to his two works, the "Tractatus Logico-Philosophicus" and "Philosophical Investigations". There's a lot explored here, but Wittgenstein basically makes the argument that the natural logic of language—how we deduce meaning from terms in a context and naturally disambiguate the semantics of ambiguous phrases—is different from the sort of formal propositional logic that forms the basis of western philosophy. However, this is also the sort of logic that allows us to apply metaphors and conceive of (possibly incoherent, possibly novel, certainly not deductively-derived) terms—counterfactuals, conditionals, subjunctive phrases, metaphors, analogies, poetic imagery, etc. LLMs have shown some affinity of the former (linguistic) type of logic with greatly reduced affinity with the latter (formal/propositional) sort of logical processing. Hallucinations as people describe them seem to be problems with not spotting "obvious" propositional incoherence.

What I'm pushing at is not that this linguistic ability naturally leads to the LLM behavior we're seeing and calling "hallucinating", just that LLMs may capture some of how humans process language, differentiate semantics, recall terms, etc, but without the mechanisms that enable rationally grappling with the resulting semantics and propositional (in)coherency that are fetched or generated.

I can't say this is very surprising—most of us seem to have thought processes that involve generating and rejecting thoughts when we e.g. "brainstorm" or engage in careful articulation that we haven't even figured out how to formally model with a chatbot capable of generating a single "thought", but I'm guessing if we want chatbots to keep their ability to generate things creatively there will always be tension with potentially generating factual claims, erm, creatively. Further evidence is anecdotal observations that some people seem to have wildly different thresholds for the propositional coherence they can spot—perhaps one might be inclined to correlate the complexity with which one can engage in spotting (in)coherence with "intelligence", if one considers that a meaningful term.

Re: Anyone got a contact at OpenAI. They have a spider problem

#167
post #116

If they don't respect robots.txt then block them using a firewall or other server config. All of these companies are parasites.

No. I've run 'bot motels myself. I've got better things to do than curating a block list when they can just switch or renumber their infrastructure. Most notably I ran a 'bot motel on a compute-intensive web app; it was cheaper to burn bandwidth (and I slow-rolled that) than CPU cycles. Poisoning the datasets was just lulz.

I block ping from virtually all of Amazon; there are a few providers out there for which I block every naked SYN coming to my environment except port 25, and a smattering I block entirely. I can't prove that the pings even come from Amazon, even if the pongs are supposed to go there (although I have my suspicions that even if the pings don't come from the host receiving the pongs the pongs are monitored by the generator of the pings).

The point I'm making is that e.g. Amazon doesn't have the right to sell access to my compute and tragedy of the commons applies, folks. I offered them a live feed of the worst offenders, but all they want is pcaps.

(I've got a list of 50 prefixes, small enough to be a separate specialty firewall table. It misses a few things and picks up some dust bunnies. But contrast that to the 8,000 prefixes they publish in that JSON file. Spoiler alert: they won't admit in that JSON file that they own the entirety of 3.0.0.0/8. I'm willing to share the list TLP:RED/YELLOW, hunt me down and introduce yourself.)

Re: Anyone got a contact at OpenAI. They have a spider problem

#168

Earlier quoted context omitted.

The end state of training on web text has always been an ouroboros - primarily because of adtech incentives to produce low quality content at scale to capture micro pennies. The irony of the whole thing is brutal.

Content you’re allowed and capable of scraping on the Internet is such a small amount of data, not sure why people are acting otherwise. Common crawl alone is only a few hundred TB, I have more content than that on a NAS sitting in my office that I built for a few grand (Granted I’m a bit of a data hoarder). The fears that we have “used all the data” are incredibly unfounded.

> Content you’re allowed and capable of scraping on the Internet is such a small amount of data, not sure why people are acting otherwise

YMMV depending on the value of "you" and your budget.

If you're Google, Amazon or even lower tier companies like Comcast, Yahoo or OpenAI, you can scape a massive amount of data (ignoring the "allowed" here, because TFA is about OpenAI disregarding robots.txt)

Re: Anyone got a contact at OpenAI. They have a spider problem

#169

Earlier quoted context omitted.

Got a source for that?

Want a real conspiracy? What do you think the NSA is storing in that datacenter in Utah? Power point presentations? All that data is going to be trained into large models. Every phone call you ever had and every email you ever wrote. They are likely pumping enormous money into it as we speak, probably with the help of OpenAI, Microsoft and friends.

> What do you think the NSA is storing in that datacenter in Utah?

A buffer with several-days-worth of the entire internet's traffic for post-hoc decryption/analysis/filtering on interesting bits. All that tapped backbone/undersea cable traffic has to be stored somewhere.

Re: Anyone got a contact at OpenAI. They have a spider problem

#170

If they follow robots.txt, OpenAI also has a bot blocking + data gathering problem too: https://x.com/AznWeng/status/1777688628308681000 11% of the top 100K websites already block their crawler, more than all their competitors (Google, FB, Anthropic, Perplexity) combined

It's not just a problem for training, but the end user, too. There are so many times that I've tried to ask a question or request a summary for a long article only to be told it can't read it itself, so you have to copy-paste the text into the chat. Given the non-binding nature of robots.txt and the way they seem comfortable with vacuuming up public data in other contexts, I'm surprised they allow it to be such an ob…

That’s the whole point. The site owner doesn’t want their information included in ChatGPT—they want you going to their website to view it instead.

It’s functioning exactly as designed.

Post reply on HN