Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

211–220 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#211

Earlier quoted context omitted.

LLMs work at all because of the high correlation between the statistically most likely response and the most reasonable answer.

That's an explanation of why their answers can be useful, but doesn't relate to their ability to "not know" an answer

FWIW, we train into kids that "I don't know" is a valid response, and when to utter it. That training is more RLHF-type than source-materal-type, too.

Re: Anyone got a contact at OpenAI. They have a spider problem

#212
post #53
post #45

Earlier quoted context omitted.

And? What are they gonna do about it (apart from making such a person/website momentarily famous).

Have you not heard of Roko's Basilisk?

Thanks for dooming everyone who reads this comment.

Re: Anyone got a contact at OpenAI. They have a spider problem

#213

Earlier quoted context omitted.

Want a real conspiracy? What do you think the NSA is storing in that datacenter in Utah? Power point presentations? All that data is going to be trained into large models. Every phone call you ever had and every email you ever wrote. They are likely pumping enormous money into it as we speak, probably with the help of OpenAI, Microsoft and friends.

As I understand it, they don't have the capability to essentially PCAP all that data.. and the data wouldn't be that useful since most interesting traffic is encrypted as well. Instead they store the metadata around the traffic. Phone number X made an outgoing call to Y @ timestamp A, call ended at timestamp B, approximate location is Z, etc. Repeat that for internet IP addresses do some analysis and then you can bui…

> most interesting traffic is encrypted as well

encrypted with an algorithm currently considered to be un-brute-forcible. If you presume we'll be able to decrypt today's encrypted transmissions in, say, 50-100 years, I'd record the encrypted transmission if I were the NSA.

Re: Anyone got a contact at OpenAI. They have a spider problem

#214

With all the news about scraping legality you'd think a multi billion dollar AI company would try to obfuscate their attempts.

If you're not walling off your content behind a login that contains terms that you agree to not scraping, then, scraping that site is 100% legal. Robots.txt isn't a legal document.

If the industry doesn't self-regulate (ie, following conventional rules and basic human courtesy) ... then it will be regulated by laws.

So let me fix what you said for you:

> Robots.txt isn't a legal document, yet.

Re: Anyone got a contact at OpenAI. They have a spider problem

#215

Earlier quoted context omitted.

Want a real conspiracy? What do you think the NSA is storing in that datacenter in Utah? Power point presentations? All that data is going to be trained into large models. Every phone call you ever had and every email you ever wrote. They are likely pumping enormous money into it as we speak, probably with the help of OpenAI, Microsoft and friends.

It would be absolutely fascinating to talk to the LLMs of the various government spy agencies around the world.

https://harpers.org/archive/2024/03/the-pentagons-silicon-va...

If they actually worked, that is.

Re: Anyone got a contact at OpenAI. They have a spider problem

#216
post #197

Earlier quoted context omitted.

Not Llama, they’ve been really clear about that. Especially with DMA cross-joining provisions and various privacy requirements it’s really hard for them, same for Google. However, Microsoft has been flying under the radar. If they gave all Hotmail and O365 data to OpenAI I’d not be surprised in the slightest.

The company that made a honeypot VPN to access competitor's traffic? They are definitively keeping their hands off their internal data, yes.

[deleted]

Re: Anyone got a contact at OpenAI. They have a spider problem

#217

Isn’t the legality of web scraping still..disputed? There’s been a few projects I’ve wanted to work on involving scraping, but the idea that the entire thing could be shut down with legal threats seems to make some of the ideas infeasible. It’s strange that OpenAI has created a ~$80B company (or whatever it is) using data gathered via scraping and as far as I’m aware there haven’t been any legal threats. Was there so…

Web scraping the public Internet is legal, at least in the U.S. hiQ's public scraping of LinkedIn was ruled to be within their rights and not a violation of the CFAA. I imagine that's why LinkedIn has almost everything behind an auth wall now. Scraping auth-walled data is different. When you sign up, you have to check "I agree to the terms," and the terms generally say, "You can't scrape us." So, you can't just make…

Most of these SaaS's have a "firehose" that if you are big enough (aka, can handle the firehose), can subscribe to. These are like RSS feeds on crack for their entire SaaS.

- https://developer.twitter.com/en/docs/twitter-api/enterprise...

- https://developer.wordpress.com/docs/firehose/

Re: Anyone got a contact at OpenAI. They have a spider problem

#218

Earlier quoted context omitted.

It's not just a problem for training, but the end user, too. There are so many times that I've tried to ask a question or request a summary for a long article only to be told it can't read it itself, so you have to copy-paste the text into the chat. Given the non-binding nature of robots.txt and the way they seem comfortable with vacuuming up public data in other contexts, I'm surprised they allow it to be such an ob…

That’s the whole point. The site owner doesn’t want their information included in ChatGPT—they want you going to their website to view it instead. It’s functioning exactly as designed.

If my web browser's extension "visits" the site and dumps it into ChatGPT for me to read its summarization of the site, what has been gained by the website operator?

Re: Anyone got a contact at OpenAI. They have a spider problem

#219
post #185

Earlier quoted context omitted.

That's an explanation of why their answers can be useful, but doesn't relate to their ability to "not know" an answer

I suspect this is going to be a disagreement on the meaning of "to know". On the same lines as why people argue if a tree falling in a wood where nobody can hear it makes sound because some people implicitly regard sound is the qualia while others regard it as the vibrations in the air.

Not really.

An LLM should have no problem replying "I don't know" if that's the most statistically likely answer to a given question, and if it's not trained against such a response.

What it fundamentally can't do is introspect and determine it doesn't have enough information to answer the question. It always has an answer. (disclaimer: I don't know jack about the actual mechanics. It's possible something could be constructed which does have that ability and still be considered an "LLM". But the ones we have now can't do that.)

Re: Anyone got a contact at OpenAI. They have a spider problem

#220

Earlier quoted context omitted.

That's an explanation of why their answers can be useful, but doesn't relate to their ability to "not know" an answer

FWIW, we train into kids that "I don't know" is a valid response, and when to utter it. That training is more RLHF-type than source-materal-type, too.

I don't follow, what does this mean to the conversation?
Post reply on HN