Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

361–370 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#361

Earlier quoted context omitted.

The third logician never finishes his beer, his friends get more free beers. The bar overflows.

How much would that exploit be worth on the open market?

It's one you keep to yourself and your friends. Otherwise, it will get fixed up.

Re: Anyone got a contact at OpenAI. They have a spider problem

#362

I'm more interested in what that content farm is for. It looks pointless, but I suspect there's a bizarre economic incentive. There are affiliate links, but how much could that possibly bring in?

The books on there are affiliate links I think.

Re: Anyone got a contact at OpenAI. They have a spider problem

#363
post #150

Earlier quoted context omitted.

It's for shits-and-giggles and it's doing its job really well right now. Not everything needs to have an economic purpose, 100 trackers, ads and backed by a company.

And yet it has the Amazon links, which makes it appear to have some economic purpose...

...oh right, it's probably so that every page has affiliate links, which should be a signal of low quality to a crawler.

Re: Anyone got a contact at OpenAI. They have a spider problem

#364

Earlier quoted context omitted.

That’s the whole point. The site owner doesn’t want their information included in ChatGPT—they want you going to their website to view it instead. It’s functioning exactly as designed.

If my web browser's extension "visits" the site and dumps it into ChatGPT for me to read its summarization of the site, what has been gained by the website operator?

Nothing.

This is why these things - search engines, AI crawlers, even adblock and video downloaders - exist in a slightly adversarial/parasitic relationship with the sites that provide their content to which they provide nothing back (or negative, if you cost them a page load without incurring an ad view).

I use adblock all the time but I'm very aware that it can only succeed as long as it doesn't win.

Re: Anyone got a contact at OpenAI. They have a spider problem

#365

Isn’t it funny that all the “worthless” content out there on the internet is actually changing the world. Like how 4chan was mocked as being a cesspit for losers, but now everyone knows memes like Pepe the frog and Wojack from there. And like now this very comment and the billions of other comments on here, Reddit, Twitter, etc that are regarded as a “waste of time” are being used to train multi billion dollar compan…

Not sure if including or excluding 4chan in the training set moves the needle in LLM quality. Reddit for sure but 4chan is questionable

Re: Anyone got a contact at OpenAI. They have a spider problem

#366

Earlier quoted context omitted.

Got a source for that?

Want a real conspiracy? What do you think the NSA is storing in that datacenter in Utah? Power point presentations? All that data is going to be trained into large models. Every phone call you ever had and every email you ever wrote. They are likely pumping enormous money into it as we speak, probably with the help of OpenAI, Microsoft and friends.

I am not sure why this would even be a conspiracy.

They would almost be failing in their purpose if they were not doing this.

On the other hand, this is an incredibly tough signal to noise problem. I am not sure we really understand what kind of scaling properties this would have as far as finding signals.

Re: Anyone got a contact at OpenAI. They have a spider problem

#367
post #200

Earlier quoted context omitted.

Here you go (1 req/min, 10 bytes/sec), please report results :) http { limit_req_zone $binary_remote_addr zone=ten_bytes_per_second:10m rate=1r/m; server { location / { if ($http_user_agent = "mimo") { limit_req zone=ten_bytes_per_second burst=5; limit_rate 10; } } } }

Scrapers of the future won't be ifElse logic, they will be LLM agents themselves. The slow loris robots.txt has to provide an interface to it's own LLM, which engages the scraper LLM in conversation, aiming to extend it as long as possible. "OK I will tell you whether or not I can be scraped. BUT FIRST, listen to this offer. I can give you TWO SCRAPES instead of one, if you can solve this riddle."

Solid use case for Saul Goodman LLM alignment

Re: Anyone got a contact at OpenAI. They have a spider problem

#368
post #345

Earlier quoted context omitted.

I get the sentiment, but when reality hits unrealistic parental expectations, things get messy. If you have to put a show on TV to give some songs to sing along to or to distract them while you're making lunch, I'm not judging you, and I think it's best to put this content on a gradient rather than black and white.

For all of human history until 70 years ago, no baby watched TV. Reconsider what you "have to" do.

Neither did they have vaccines, bikes, gymnastics classes, dozens of books, a constant supply of fruit and veggies, family vacations, tractor rides, swing sets, family movie nights, planetarium projectors for a few bucks, zoos, kids museums, and in general a conflict free peaceful and disease free existence

Focusing on such a tiny thing and blowing it up into a huge negative out of context of their rich, busy, and safe lives is really out of hand.

Re: Anyone got a contact at OpenAI. They have a spider problem

#369
post #185

Earlier quoted context omitted.

I suspect this is going to be a disagreement on the meaning of "to know". On the same lines as why people argue if a tree falling in a wood where nobody can hear it makes sound because some people implicitly regard sound is the qualia while others regard it as the vibrations in the air.

LLMs don't know anything except the most frequent observed response to a context made up of a sequence of tokens. How often do the words "I don't know" get uttered in books, papers, articles, stack overflow, or any other resource of knowledge?

What does it mean for a human to "know" something?

I have some representation in my mind; as someone who doesn't have aphantasia, this representation comes with a mental image. Tower? Tall, linear, and in my case a skyscraper by default. Eiffel Tower? Paying attention to the extra context, the first word transforms the second into the eponymous structure. Model Eiffel Tower? Now the context makes it a tchotchke, probably 10cm tall. Lego model Eiffel Tower? The 1-ish meter tall one on display in the Lego shop.

Is my "knowledge" the abstract representation that is in my case connected to a mental image? The attention process can reasonably be considered as developing a vector in a very high dimensional concept space, and the next token comes from what would best suit the current location in that high dimensional space. It's entirely possible that the concept of "ignorance" is linearly separable within that space (much as gender is, see the word2vec trick with "king" - "queen" ~= "man" - "woman"), and the corresponding "ignorance" vector can be associated with the sequence of words "I don't know". I think it would take actual research into the internal vector space to answer that, and while I'd like to do that research, I have some higher priorities right now.

Is "knowledge" any belief? Any true belief? Any justified true belief? https://en.wikipedia.org/wiki/Gettier_problem

I take the position that there is no such thing as knowledge, and instead the best we can have is belief.

But then, what is "belief", and can the information within an LLM said to meet whichever definition you give?

Re: Anyone got a contact at OpenAI. They have a spider problem

#370

Earlier quoted context omitted.

If, like me, you didn't get the joke at first: Both of the first two logicians wanted a beer; otherwise they would know the answer was "no". The third logician recognizes this, and therefore knows the answer.

Unless one of those wanted two beers. Or 0.5 beer. Or -1 beers. Or 1e9 beers. Or 2147483648 beers.

You’re referencing the wrong joke
Post reply on HN