Live data from Hacker News

June 2023 Data Dump is missing

meta.stackexchange.com

171–180 of 268 posts

Re: June 2023 Data Dump is missing

#171
post #117

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Network_News_Transfer_Protocol Originally over UUCP (Unix Unix Copy Protocol) and done via dial ups at night (when the rest of the batch transfers were done - email too with the old bang path). The two servers would exchange all the batched email and news posts that were routed to the other side. RFC 977 ( https://www.w3.org/Protocols/rfc977/rfc977 ) has an example of how files are copie…

But how did all of this line up within the federated or not conversation? If each ISP could host their own version, that doesn't sound federated. But who was in control of the "main" source of truth type of version?

There was no "main" source of truth for each version and each ISP could have its own set of posts. The bofh.* hierarchy for example had a very limited distribution and if it was found that one of the sites that provided it was leaking it to the general public they'd collectively cut off sites from being able to post or receive posts until they rectify their configuration.

http://usenet.trigofacile.com/hierarchies/index.py?status=pr... and https://web.archive.org/web/20220815151921/http://bofh.taron...

Sure, an ISP could host its own news.answers site and not accept posts from others for that group nor pass messages from it out to others... but that would be a rather lonely place.

Likewise, one ISP may have a different set of posts that it shows for a news group because of moderation actions (we don't accept any posts greater than 1MB in size because of disk space issues and we don't host any newsgroup named *.binaries).

Federation comes from exchanging those posts using a common protocol - NNTP.

Re: June 2023 Data Dump is missing

#172
post #117

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Network_News_Transfer_Protocol Originally over UUCP (Unix Unix Copy Protocol) and done via dial ups at night (when the rest of the batch transfers were done - email too with the old bang path). The two servers would exchange all the batched email and news posts that were routed to the other side. RFC 977 ( https://www.w3.org/Protocols/rfc977/rfc977 ) has an example of how files are copie…

But how did all of this line up within the federated or not conversation? If each ISP could host their own version, that doesn't sound federated. But who was in control of the "main" source of truth type of version?

Because there is no central controller of NTTP feeds or data streams.

Re: June 2023 Data Dump is missing

#173

What irks me about this is that 100% of their data is provided for free, by the community that they have fostered, the people like myself who have answered > 2500 questions[0], and now SO feels hard-done-by by LLMs using all their hard work to create tools like CodeGPT, GitHub copilot, etc. Were it really a site for helping developers to improve their skills and increase their productivity through the give-and-take m…

Fair enough, but you were fully aware of this arrangement going in, and chose to participate. SO didn't opt into being training data for ChatGPT, and I doubt they would have given the chance. You may object that SO implicitly did so by making their site available to the public, but the ethics of GAI training data a new moral gray area that we're still navigating. They at least have something of a case to be made.

Re: June 2023 Data Dump is missing

#174
post #41

Earlier quoted context omitted.

Nitpick, also because the contrast is kind of funny: mote: a small particle, speck, atom, "mote of dust" moat: a deep ditch, often filled with water, as a first line of defence around a castle.

Hopefully this comment won't be demoated by the algorithm - it truly holds water on its own!

You are only helping it with puns and we know that puns are the gateway to consciousness! They are like the corpus callosum of language, serving as the bridge that spans the moat between the castles of wit and creativity.

Re: June 2023 Data Dump is missing

#175

What irks me about this is that 100% of their data is provided for free, by the community that they have fostered, the people like myself who have answered > 2500 questions[0], and now SO feels hard-done-by by LLMs using all their hard work to create tools like CodeGPT, GitHub copilot, etc. Were it really a site for helping developers to improve their skills and increase their productivity through the give-and-take m…

Fair enough, but you were fully aware of this arrangement going in, and chose to participate. SO didn't opt into being training data for ChatGPT, and I doubt they would have given the chance. You may object that SO implicitly did so by making their site available to the public, but the ethics of GAI training data a new moral gray area that we're still navigating. They at least have something of a case to be made.

> Fair enough, but you were fully aware of this arrangement going in, and chose to participate.

Yep, but that's not the point.

> SO didn't opt into being training data for ChatGPT, and I doubt they would have given the chance.

Neither did Wikipedia (at least to my knowledge). I thought the point of opening up information was to benefit the public, first and foremost, and without hidden terms which state something along the lines of "it's free and open information built by the community, but when something disrupts our ads-driven business model and we make it unfree".

It would have been nice if they had at least allowed their contributors to vote on this, or have some sort of a say.

Re: June 2023 Data Dump is missing

#176

"Just sorta stating the obvious here, but the timing of this is unbelievably terrible; I actually can't fathom a worse time for this call to be made than in light of this week. –zcoop98" Or, it's exactly the best time to do it. Doing it now allows your news to get blended in with the Reddit news. Doing it later after Reddit chatter settles down means all of the chatter is directed squarely at you.

It also means people are more motivated to build a replacement than just by the timewasting reddit being unavailable.

Replace them both with a model somewhat like Wikipedia, open the content for the world, and get a cut of profits from the corporations that want to use the data to train on it.

Re: June 2023 Data Dump is missing

#177
post #145

Earlier quoted context omitted.

It’s a fork of the community rather than the data and content.

You really would need an existing dump to seed the new site.

There were a number of issues that lead to the decision not to do a grab and seed of SE into codidact.

There was the "what license is that post actually under? Is it 2.5? 3.0? 4.0?" which made things difficult.

There was the "what are the actual attribution requirements that SE has for sites that use its content?" This is a bit of an issue because it's never really clear what those requirements are and what you need to do. It can also hurt SEO because it's duplicated content. Furthermore, codidact leadership had already and enough dealings with SE lawyers and likely wanted to avoid any other.

Lastly, there was the desire to make a philosophical break with SE. The codidact founders didn't want to have anything to do with SE.

Some sites are doing ok. Others stood up but didn't have sufficient involvement to keep them going.

For a counter example, "PhysisOverflow" has an import tool that they use. https://www.physicsoverflow.org/4536/import-queue

Having an imported site that is mostly inactive with activity on that same content is even more disappointing than having a mostly empty site. And active mirroring is a time-consuming process that runs into rate limit issues with an API.

Re: June 2023 Data Dump is missing

#178

Earlier quoted context omitted.

This was one of the promises originally of Stack Overflow: all the content is Creative Commons licensed so that if they "turned evil" (I believe it was Joel that put it this way) the community could, in a way, create a fork. https://web.archive.org/web/20230203170609/https://stackover... Unfortunately the dumps themselves are not a legal requirement, just a gentleman's agreement, so realistically exercising this abil…

I always wonder why original founders just sell the company and do something else. Why don't they try to control it more and make sure it stays aligned with needs of society more? Either they can't because of shareholder/equity owners pressure, or they won't, because they really don't care and just said it for PR

> Either they can't because of shareholder/equity owners pressure, or they won't, because they really don't care and just said it for PR

That is assuming the worst in people. Have you ever wanted to move onto something new? If you make something cool, it is not your lifelong obligation to oversee it.

Re: June 2023 Data Dump is missing

#179

Earlier quoted context omitted.

I always wonder why original founders just sell the company and do something else. Why don't they try to control it more and make sure it stays aligned with needs of society more? Either they can't because of shareholder/equity owners pressure, or they won't, because they really don't care and just said it for PR

The original founders sold the site for $1.8 Billion.

That was in 2021. https://stackoverflow.blog/2021/06/02/prosus-acquires-stack-...

Joel left in 2019 https://twitter.com/spolsky/status/1111267189133316097

Jeff left in 2012 https://blog.codinghorror.com/farewell-stack-exchange/

I'm not sure that it is fair to say that the original founders sold it for that amount.

Re: June 2023 Data Dump is missing

#180
post #125

Earlier quoted context omitted.

Sounds like you need to allowlist google from your grey list. They have a long retry on their side. Once you tell a server to 'go away, come back later' it is really up to the server to decide when or if to retry. Additionally, If they use multiple sending IPs, you can end up grey listing again and again before they try back with a good ip. You'd either need to allowlist the big providers sending blocks or just drop…

I completely disabled greylisting to try and resolve that, but it is another good point on why gmail kind of sucks for people not on gmail.

So you accepted their mail on first contact? And first contact was 20 min after sent header timestamp?
Post reply on HN