I can't believe nobody has mentioned naive Bayesian text classification yet. It sounds like it could work wonders for Twitter. I'm much more likely to be interested in tweets with words like "hylomorphism" than tweets with words like "omglol", and a text classification algorithm could learn that if you trained it up some. It doesn't have to be perfect; it just has to improve the signal-to-noise ratio significantly.
I've looked into this a bit, albeit more in a spam-filtering context; tweets have very little text for naive Bayes to latch onto. 140 characters would be 20-30 words, tops. That is so few words that it is hard to move the prior very much, unless there are blockbuster words that almost always indicate a bad tweet; as the article suggested, "breakfast", "beer", etc.
Twitter's garbage problem
31–40 of 62 posts
Re: Twitter's garbage problem
#32I've never met the author of the article, but I assume he's the type of person who is bothered when his Reader unread tally switches over to the plus mark. I mean, honestly, tweets are 140 characters or less. The average tweet takes a handful of seconds to read. Is my time so important that I can't spend a few minutes of my day learning what my friends thought was important to share with me? Must I tailor their inter…
Re: Twitter's garbage problem
#33I can't believe nobody has mentioned naive Bayesian text classification yet. It sounds like it could work wonders for Twitter. I'm much more likely to be interested in tweets with words like "hylomorphism" than tweets with words like "omglol", and a text classification algorithm could learn that if you trained it up some. It doesn't have to be perfect; it just has to improve the signal-to-noise ratio significantly.
I've looked into this a bit, albeit more in a spam-filtering context; tweets have very little text for naive Bayes to latch onto. 140 characters would be 20-30 words, tops. That is so few words that it is hard to move the prior very much, unless there are blockbuster words that almost always indicate a bad tweet; as the article suggested, "breakfast", "beer", etc.
Re: Twitter's garbage problem
#34Earlier quoted context omitted.
I've looked into this a bit, albeit more in a spam-filtering context; tweets have very little text for naive Bayes to latch onto. 140 characters would be 20-30 words, tops. That is so few words that it is hard to move the prior very much, unless there are blockbuster words that almost always indicate a bad tweet; as the article suggested, "breakfast", "beer", etc.
Dont many classifiers pick the most interesting n words and only use those to decide? Where n is something like 15.
Re: Twitter's garbage problem
#35I can't believe nobody has mentioned naive Bayesian text classification yet. It sounds like it could work wonders for Twitter. I'm much more likely to be interested in tweets with words like "hylomorphism" than tweets with words like "omglol", and a text classification algorithm could learn that if you trained it up some. It doesn't have to be perfect; it just has to improve the signal-to-noise ratio significantly.
Re: Twitter's garbage problem
#36Earlier quoted context omitted.
I disagree. Some of the tweets I subjectively label as garbage might be legit. Hey, maybe some guy's mom is on twitter and wants to know what he had for lunch. It's not a people problem, it's a drawback of the platform that all of this different subject matter has to be broadcast in the same stream.
It's a people problem, but the problem is you. If you don't care about what he's having for lunch, then I'm not so sure what's so hard about ignoring his tweet. It's only going to be on your screen for a split second as you're scrolling through other stuff. Presumably anyone you follow would post more things that are worthwhile than not; otherwise, you should just find new people to follow. One of my friends tweets q…
Re: Twitter's garbage problem
#37I've never met the author of the article, but I assume he's the type of person who is bothered when his Reader unread tally switches over to the plus mark. I mean, honestly, tweets are 140 characters or less. The average tweet takes a handful of seconds to read. Is my time so important that I can't spend a few minutes of my day learning what my friends thought was important to share with me? Must I tailor their inter…
I respect your sympathy for and willingness to read what those you follow deem interesting. But I think the author of the blog post is speaking of a problem that emerges at a larger scale when you follow more & more people. Your stream becomes bigger while the amount of time you have stays the same so your options are 1) unfollow people, 2) don't read everything, or 3) put in more time. Personally, I don't like 1 or…
I just feel like I personally get information in so many different ways, that if something is big enough, I'll see it in many places and am pretty much bound to read one of them. I can count on seeing any big tech story here on five different feeds, Google News (go to time-killer on my phone in the rare occasion my feeds are empty), several twitter accounts, and probably here on HN (more like 15 feeds if it has Apple in the title.)
I went to your site but you lost me with the lack of content--maybe at this stage that's important to weed out the laziest of the "testers." Have you thought of doing just a quick before/after screenshot to show default twitter vs. your application, or is it too early for that?
Edited to add: er, my fault. My brain skips over flash almost automatically. I didn't watch your video.
Re: Twitter's garbage problem
#38I've always thought of this problem in reverse for both Twitter and Facebook. I have certain followers/friends that are interested in my thoughts about programming and business and others who would care more about where I'm going this afternoon. It would be nice if there were different publishing channels I could publish to different friends.
You can do that on Facebook: when posting a status update, the drop-down box under the lock icon has a "Customize" option, which will pop up a dialog where you can select specific people or pre-created lists of friends for it to be visible to (the lists function essentially as channels, like "colleagues" or "family"). You can also set one of those channels as the default. I know a few people who do that fairly regula…
Re: Twitter's garbage problem
#39I've always thought of this problem in reverse for both Twitter and Facebook. I have certain followers/friends that are interested in my thoughts about programming and business and others who would care more about where I'm going this afternoon. It would be nice if there were different publishing channels I could publish to different friends.
You can do that on Facebook: when posting a status update, the drop-down box under the lock icon has a "Customize" option, which will pop up a dialog where you can select specific people or pre-created lists of friends for it to be visible to (the lists function essentially as channels, like "colleagues" or "family"). You can also set one of those channels as the default. I know a few people who do that fairly regula…
Re: Twitter's garbage problem
#40I honestly don't see this as a problem. I don't read every Tweet that comes through my stream its kind of random access information, I look at twitter every now and then and if something interesting strikes me I look into it. While filtering would be a good feature, I think saying its a killer feature is going a bit far. I don't think the reasoning that you could follow twice as many people would make much sense from…
Custom searches don't capture all of the good information that's out there because I may not specify the right keywords. It's ANTI-search that I'm personally looking for.