Live data from Hacker News

June 2023 Data Dump is missing

meta.stackexchange.com

201–210 of 268 posts

Re: June 2023 Data Dump is missing

#201
post #16

This, along with recent Reddit goings-on has made me realize a major risk with the current structure of online communication. Take either Reddit or Stack Exchange as examples. They build a platform, and users contribute their time, thought, energy, and knowledge to build a community on that platform. Those companies can then gatekeep and restrict access to all that the community built, when all they did is provide th…

>We need to rethink this model.

Once upon a time, most people who wanted or had something to say wrote their own little website and hosted it themselves (be it in a datacenter or a server in their closet). Some even ran forums and got fancy with server-side magic because that's what nerds do. Even the kids who couldn't afford anything had free, basic hosting services to choose from (anyone remember those days?).

The internet was designed as a distributed network and the denizens then were distributed. You only got as centralized as a given ISP or datacenter provider.

Of course, we all know as more and more commoners came onto the internet they didn't want to bother with developing or hosting or maintaining a website or anything. They just wanted to shitpost, for free, with blackjack and hookers.

And so "free" services like Reddit, Facebook, et al. came about to serve that demand. Information became centralized, because who the fuck has time to be responsible? Offload that crap!

The cost of that offloading of responsibility has now come knocking with debt collectors in tow, with interest.

I guess what I'm trying to say is: We don't need to rethink anything. We just need to take some god damn responsibility for ourselves. Responsibility is power, and with power you can tell commercial interests you disagree with to screw off.

Re: June 2023 Data Dump is missing

#202
post #140

Earlier quoted context omitted.

I always wonder why original founders just sell the company and do something else. Why don't they try to control it more and make sure it stays aligned with needs of society more? Either they can't because of shareholder/equity owners pressure, or they won't, because they really don't care and just said it for PR

Because despite claims to the contrary most of these sites/projects aren't created for altruistic reasons, they were created to make money (at some point). Cashing out is typically part of the long term plan. In the case of Stack Overflow, I think the reason for the data dumps was two-fold: one of the original founders (who left long ago) came across as at least idealistic and wanting to do the right thing. The other…

Wow I read the text for that link you posted in a very different way than I intended.

Re: June 2023 Data Dump is missing

#203
post #63

Earlier quoted context omitted.

Time to adopt Nostr as future-proof path.

Without movement on this [1] I can't see adoption. [1] https://github.com/nostr-protocol/nostr/issues/97

So, so much decentralized tech never gets adoption due to a lack of an identity management layer that nobody wants to build because it can’t be perfectly decentralized and have the account recovery features that 99% of regular folks need. This is an example where perfect is the enemy, nemesis even, of good.

Someone should build an identity system that is optionally centralized or federated (if you like your key custody, you can keep it), migrateable and that ONLY handles identity. That will still be orders of magnitude better than relying on Google, Twitter and friends, simply because there won’t be a glaring conflict of interest of platform rent-seeking.

Moreover, anyone who wants to build decentralized/federated apps don’t have to reinvent the wheel poorly. It’s so sad to see project after project fading into the ether because people can’t fucking sign in in a reasonable way.

At least with crypto currency, there’s a somewhat strong argument for individual key custody, but I’m not talking about protecting $20M while on the run from the feds, I’m talking about afternoon shitposting with friends and strangers.

Re: June 2023 Data Dump is missing

#204
post #16

This, along with recent Reddit goings-on has made me realize a major risk with the current structure of online communication. Take either Reddit or Stack Exchange as examples. They build a platform, and users contribute their time, thought, energy, and knowledge to build a community on that platform. Those companies can then gatekeep and restrict access to all that the community built, when all they did is provide th…

This was one of the promises originally of Stack Overflow: all the content is Creative Commons licensed so that if they "turned evil" (I believe it was Joel that put it this way) the community could, in a way, create a fork. https://web.archive.org/web/20230203170609/https://stackover... Unfortunately the dumps themselves are not a legal requirement, just a gentleman's agreement, so realistically exercising this abil…

Maybe just a gentlemen's agreement, but a nice canary too. Once the dumps stop, it's time to start waving middle fingers and GTFO.

Re: June 2023 Data Dump is missing

#205

Earlier quoted context omitted.

This was one of the promises originally of Stack Overflow: all the content is Creative Commons licensed so that if they "turned evil" (I believe it was Joel that put it this way) the community could, in a way, create a fork. https://web.archive.org/web/20230203170609/https://stackover... Unfortunately the dumps themselves are not a legal requirement, just a gentleman's agreement, so realistically exercising this abil…

So the idea is that in case leadership wants to 'carve out a kingdom' that is not in line with community wishes, the community could take the data dump and create a clone of sorts? Then now the last snapshot for doing so would be the last data drop from March?

Yes. There's moderately successful precedent: Wikivoyage is a fork of Wikitravel, which was went evil after it was sold to a content farm.

Re: June 2023 Data Dump is missing

#206
post #195

What irks me about this is that 100% of their data is provided for free, by the community that they have fostered, the people like myself who have answered > 2500 questions[0], and now SO feels hard-done-by by LLMs using all their hard work to create tools like CodeGPT, GitHub copilot, etc. Were it really a site for helping developers to improve their skills and increase their productivity through the give-and-take m…

> None of the contributors (apart from the employee ones, I suppose) ever got paid any currency other than high-fives in the form of rep, medals, the gamified stuff, moderation rights, and at certain rep levels some swag in the form of t-shirts and the usual. I would love to see some kind of identity and reputation system where the "high-fives in the form of rep" could follow people across communities. It may not fee…

> I would love to see some kind of identity and reputation system where the "high-fives in the form of rep" could follow people across communities. It may not feel like much compensation if you've contributed over 2500 answers, but having reputation gained in your area of expertise grant you a high level of trust to interact in other communities could be valuable, at least in my opinion.

Honestly I think that's an excellent idea - a rep "passport" of sorts which gains you a certain level of trust within certain communities.

> Assuming they're making this move to protect against AI / LLMs, I think SO is in an impossible situation here. When all the ChatGPT hype started, one of my first questions was "what happens to the incentive for contributors and creators?" Why would I want to contribute on a platform if I know an AI model is going to come in, take my contribution, and regurgitate it back to the masses in a way that I can't control?

Sadly, I think this is an unpreventable outcome of what is happening right now. I don't think anyone will have any control over this, at all. We can only hope it will never be the case that being active (actual human contributors) becomes a worthless pursuit.

> Even if I get some attribution from the AI/LLM, do I even want it? If the LLM is blending content from multiple sources, which changes the context and presentation I put effort into, is the quality going to be high enough to match what I strive to achieve for myself when I'm trying to build a reputation as a high quality contributor? What if the AI is hallucinating objectively poor quality content and giving me partial attribution?

Another excellent point, the prospect of this being possible today - AI attribution from a hallucinated version of a human's objective contribution sounds freaking terrifying to me. Not a world I want to live in, to be honest.

> I think AI is going to be disruptive and the whole idea, for me anyway, behind disruption is that you break an existing system and then everyone is free to take a shot at claiming part of the new gold rush that occurs while trying to build the replacement. The problem with AI is that it's going to break a lot of services that do a good job of serving the community and shouldn't be broken. SO is a great example of a healthy community that doesn't need disruption, but the massive amount of high quality, curated content is going to make them a prime target for LLM training.

As will every single human-created/curated content-source, IMHO. I think that "quality" will be really, really hard to objectively measure in the near future as the whole world of digital information becomes tainted with applied statistical models which can do a reasonably good job of predicting what people perceive to be high-quality reasoning, answers, content. I like the idea of underground speakeasies where there's no wifi, just humans.

> Personally I think the only solution is for "noai" variants of popular open source licenses so contributors have the ability to make it clear they don't want to contribute to AI/LLM companies. If SO had an option to flag contributions as CC-BY-SA-NOAI, I'd enable it on my stuff going forward.

That would be great, but I'm pretty sure that no LLM corporation would care about those flags, even with strict regulations in place from governments.

Re: June 2023 Data Dump is missing

#207
post #16

This, along with recent Reddit goings-on has made me realize a major risk with the current structure of online communication. Take either Reddit or Stack Exchange as examples. They build a platform, and users contribute their time, thought, energy, and knowledge to build a community on that platform. Those companies can then gatekeep and restrict access to all that the community built, when all they did is provide th…

> We need to rethink this model.

This problem is inherent to client/server software, and there are really only three ways to do it:

1. The server side of client/server is centralized and run by corporations

2. The server side is decentralized, meaning everyone has their own server

3. Abandon the server, clients connect directly to each other without a server intermediating

Option 3 would be ideal, but would require significant technological advances - it'll be a lo0ong time before bandwidth is cheap enough that Kim Kardashian can serve photos and movies to all of her fans direct from her phone. Option 1 is what we have now, and is terrible in a variety of ways.

Option 2 would be hard but is not obviously impossible, so still our best bet - sure, it's not viable now, but it sure seems like it could be, if an iphone's worth of r&d were put in to it. I would honestly be amazed if no one at Amazon is working on such a thing, since no one would benefit more than AWS from a future in which a cloud VM becomes one of the things that most middle-class families rent monthly.

Re: June 2023 Data Dump is missing

#208
post #16

This, along with recent Reddit goings-on has made me realize a major risk with the current structure of online communication. Take either Reddit or Stack Exchange as examples. They build a platform, and users contribute their time, thought, energy, and knowledge to build a community on that platform. Those companies can then gatekeep and restrict access to all that the community built, when all they did is provide th…

> We need to rethink this model. This problem is inherent to client/server software, and there are really only three ways to do it: 1. The server side of client/server is centralized and run by corporations 2. The server side is decentralized, meaning everyone has their own server 3. Abandon the server, clients connect directly to each other without a server intermediating Option 3 would be ideal, but would require s…

Content-addressing together with P2P and extra paid relays for those who really need it. In terms of "superstars" sharing content, if they share their image which is content-addressed and can be fetched from anyone, it's enough that one peer shares it with three others for it to be reliable enough in practice. Content like that is also usually just relevant because of recency, so large swaths of people try to access it within 24h, after that the news cycle already moved on so won't be fetched much after that.

Re: June 2023 Data Dump is missing

#210
post #179

Earlier quoted context omitted.

The original founders sold the site for $1.8 Billion.

That was in 2021. https://stackoverflow.blog/2021/06/02/prosus-acquires-stack-... Joel left in 2019 https://twitter.com/spolsky/status/1111267189133316097 Jeff left in 2012 https://blog.codinghorror.com/farewell-stack-exchange/ I'm not sure that it is fair to say that the original founders sold it for that amount.

It's possible that they had equity in the company after they left.
Post reply on HN