Live data from Hacker News

Twitter now requires an account to view tweets

techcrunch.com

941–950 of 1001 posts

Re: Twitter now requires an account to view tweets

#941
post #650

Earlier quoted context omitted.

HN isn't next. We hate change. (More precisely: we're acutely aware that users hate change, and since we do too, it's kind of an easy call.) Also, the value of HN to YC consists of the community and keeping the community happy is therefore a must. (Did I say happy? More precisely: as happy as possible under the circumstances)

This attitude that: if it ain't broke, don't destroy it, is something I'm finding myself valuing increasingly. I use Stylish to make HN look a bit more readable and prettier, and beyond that, it functions exactly how I want it to and I'm glad to see that the institutional momentum here is a core value.

Take a look at my profile and CSS hacks, if you're interested.

(I've a slightly more updated set locally, can share those if requested.)

Re: Twitter now requires an account to view tweets

#942
post #76

Paul Graham famously started to use Mastodon (but has not written anything there since last year). But the HN Status emergency “is HN down” channel is still only on Twitter. It used to be publicly readable at https://twitter.com/HNStatus >. But now, if HN was to go down, only logged-in Twitter users would be able to see why.

The link from robotblake's redirector rules works (for now):

https://syndication.twitter.com/srv/timeline-profile/screen-...

https://news.ycombinator.com/item?id=36544185

Re: Twitter now requires an account to view tweets

#943
post #783

Earlier quoted context omitted.

The choice is not between no change and drastic change, but selecting a rate of change that is appropriate. Things which do not change, die, as the only thing that is constant is change. Change or be changed.

Sharks haven't changed in about 450 million years. There are designs that just work and don't need to change unless the environment changes drastically.

Well, actually...

https://www.smithsonianmag.com/smart-news/reef-sharks-are-di...

https://www.fishersci.com/us/en/scientific-products/publicat...

https://www.science.org/content/article/great-white-sharks-h...

Re: Twitter now requires an account to view tweets

#945
post #580

Elon Musk wrote: Several hundred organizations (maybe more) were scraping Twitter data extremely aggressively, to the point where it was affecting the real user experience. What should we do to stop that? I’m open to ideas. https://twitter.com/elonmusk/status/1674887204580073474

They can still scrap it but now with an account cookie? Also, what’s the problem of scraping it, it’s just like browsing systematically.

If there is a sufficient quantity of scrapers, it's equivalent to a DOS attack. The scraping traffic may even exceed the user traffic.

Re: Twitter now requires an account to view tweets

#946

From Elon's twitter ( https://twitter.com/elonmusk/status/1674942336583757825?s=20 ) "This will be unlocked shortly. Per my earlier post, drastic & immediate action was necessary due to EXTREME levels of data scraping. Almost every company doing AI, from startups to some of the biggest corporations on Earth, was scraping vast amounts of data. It is rather galling to have to bring large numbers of servers online on an…

Oh how the disruptors hate disruptors

Re: Twitter now requires an account to view tweets

#947
post #650

Earlier quoted context omitted.

HN isn't next. We hate change. (More precisely: we're acutely aware that users hate change, and since we do too, it's kind of an easy call.) Also, the value of HN to YC consists of the community and keeping the community happy is therefore a must. (Did I say happy? More precisely: as happy as possible under the circumstances)

It's ambiguous what GP refers to as "next", but if it's the Eternal September part, I believe HN unfortunately already suffers a lot from it. In my subjective opinion, comment quality took a nosedive in the past 2-3 months or so. That said, I have no idea how to fix it, if it needs fixing at all.

Paralleling what dang's said here, I've been looking at 17 years of HN front page activity over the month or so, and am starting to tackle the question of topic drift and/or focus over that period.

I've been sort of live-blogging the experience on the Fediverse: https://toot.cat/@dredmorbius/tagged/HackerNewsAnalytics>, as well as in some of my HN posts.

My current tack involves looking at sites (as reported in parentheses at the end of each HN front-page post title) and classifying those. With slightly more than 30% of sites categorised, I can classify about 65% of all HN posts.

For the full dataset (17 years), that's roughly:

     1  63913  35.73%  UNCLASSIFIED
     2  22589  12.63%  blog
     3  15112   8.45%  general news
     4  13823   7.73%  tech news
     5  12851   7.18%  programming
     6   8622   4.82%  corporate comm.
     7   8459   4.73%  academic / science
     8   7294   4.08%  n/a
     9   5324   2.98%  business news
    10   3803   2.13%  general interest
    11   2151   1.20%  social media
    12   2074   1.16%  software
    13   1613   0.90%  technology
    14   1463   0.82%  video
    15   1144   0.64%  general info (wiki)
    16   1009   0.56%  government
    17    724   0.40%  misc documents
    18    720   0.40%  law
    19    702   0.39%  tech discussion
    20    620   0.35%  science news
Tons of caveats: this depends heavily on how I classify individual sites, a given site's stories might well be technical, social, or political, etc., etc.

The breakdown-by-year analysis is in development, but if anything programming-specific content as increased in prevalence. Political discussion seems not to have (though it rose significantly ~2014). Cryptocurrency and blockchain-specific sites also peaked about that time (I suspect much of that discussion is now mainstream). General news has always been a huge portion of HN discussion, as have individual (and corporate) blogs.

Note again that this isn't about discussion and comments, or even the titles or article contents (I'm thinking of looking at those, it's ... a challenge for me).

But across nearly 200,000 front-page stories, on which nearly half of all HN discussion occurs (based on another API-based study looking at comprehensive posts), the overall trending seems at first blush to be pretty consistent and if anything improving over time.

(As with all preliminary results, I'm hoping I won't have to eat my words here. Though I'm reasonably confident in most of this.)

From the classifications above, the places you might find some that "suffering" would be in general news, genral interest, and social media categories. All but the first of those are single-digit percentages, and a lot of that general-news content is about technology, business, finance, and science, all of which would crowd out the sort of social and political issues which seem to generate strong feelings.

The "UNCLASSIFIED" sites are a wide mix, though most are probably a mix of blogs, corporate / organisational communications, and the like. The mean posts per site is 1.739951, so gains from additional site-categorisation are pretty slim. I have captured a lot of obvious patterns via regexes and string matches, so academic/science and major (or even minor) blogging and social media sites aren't a large fraction.

More recent discussion here: https://news.ycombinator.com/item?id=36524001>

Re: Twitter now requires an account to view tweets

#948
post #871

Earlier quoted context omitted.

It's ambiguous what GP refers to as "next", but if it's the Eternal September part, I believe HN unfortunately already suffers a lot from it. In my subjective opinion, comment quality took a nosedive in the past 2-3 months or so. That said, I have no idea how to fix it, if it needs fixing at all.

I'm not saying you're wrong and anyway it is hard, if not impossible, to evaluate objectively—but I can tell you two things for sure. One is that people have been saying more or less exactly this about HN for at least 15 years; the other is that HN is subject to a lot of random fluctuations, and random swings tend to get interpreted by humans as long-term trends—not because they are, but because that is what humans d…

In addition to my sibling comment: HN also steps in to quash developing negative patterns in all sorts of ways. There's a long list of banned sites, there is the flamewar detector (though I've ... questions ... about that), dupes detection (or flagging), there are weightings and penalties given to various sites. I believe also some keyword and other patterns are looked for as well, "Reddit" being among ones dang's recently discussed.

So yes, occasionally some new pattern or trend will emerge, but HN adapts to those fairly quickly.

Re: Twitter now requires an account to view tweets

#949

Earlier quoted context omitted.

I appreciate that attitude and value it myself, but I like to point out that it is not without risk. If the world around HN (including its community) changes, stasis can damage or kill it as well. Specifically regarding the issue of the original posting: - HN is already an important data source for large language model training. [1] - To the best of my knowledge there is no freely downloadable and current data dump o…

I don't expect HN give a fuck about the scraping. It's pure HTML, no images, probably cached all to hell for users who aren't logged in anyway. The one thing I see as a future issue is that people are starting to post comments that clearly look like they were manufactured by ChatGPT and friends. Or that could just be the way some people talk and I've spent too long with ChatGPT now and start to smell it everywhere.

HN does have performance / capacity issues, and you'll find that if you're crawling the site rapidly, you'll quickly have your IP banned.

I've had that happen even under manual browsing (when logged out). My front-page analytics project hit that limit quickly (within about 30 requests, probably less). Adding in a reasonable delay got around that.

Keep in mind that a lot of Web infrastructure tends over time to operate just at the edge of stability, as capacity costs money.

Re: Twitter now requires an account to view tweets

#950

Earlier quoted context omitted.

Elon has a good point there. Much of the current AI hotness is predicated on stealing peoples content and exploiting the infrastructure that other people have built. I don’t think it’s acceptable. The licenses, compensation models, law, technical solutions, attribution, security and privacy all need time to catch up. Regulation has a role to play as its a bit of a free for all right now. The irony of Elon mentioning…

Please don't throw around the word "stealing" so loosely. Scraping data from a public website is not "stealing". It might be a violation of the terms of service, but then you have the whole issue of click-through (formerly shrink-wrap) licenses and contracts of adhesion. If someone isn't vetting you and potentially signing you to a more meaningful contract before giving you access, for free, to data, then using that…

There are much bigger issues.

1. Users or one might say content creators don't own their data. Not just do those platform owners make a lot of money with the content (which they have a license to, as per site ToS) but then you have third parties scraping it now for commercial products. Using the data to train models that are then sold back to some of the same social media users who produced the content for free in the first place wasn't a thing until very recently, it used to be a select few doing machine learning research in the past. The laws are lacking behind the tech development and regular internet users are being exploited because of it.

2. It absolutely is stealing in some cases, and even worse. For example when they scrape it for content which they then use to train their bots to impersonate humans. Or on Twitter, there's a very common type of bot that steals content from young attractive female social media users in China, auto-translated to English, to pose as them. If you're in finance and crypto circles they're swarming with these accounts (guess the scammers know their targets).

3. In general this is only going to get worse from here on. LLM are getting better and better. On sites like Twitter you already have no idea if you're interacting with a human or not. But these "AI" can not actually think for themselves, they can only emulate, they can copy other humans. At least so far. So for the sake of making progress and ensuring we can still have intelligent discussions and find novel ideas online, it's imperative to have a way to keep the machines out. Social media must become sybil resistant or it dies in a vicious circle of self-referencing bots ever parroting the same old talking points, or variations thereof. We urgently need human ID!

Post reply on HN