Live data from Hacker News

Deep-Fried Data

idlewords.com

71–80 of 149 posts

Re: Deep-Fried Data

#71

Earlier quoted context omitted.

grad students are put in the same category as interns and teenagers, a naive type of person still in the making. i dont think there's any ill will intended.

You are dim witted. No ill will intended.

There's a difference between metaphors/jokes and insults addressed at you personally.

Re: Deep-Fried Data

#72
post #27
post #3

"The names keep changing—it used to be unsupervised learning, now it’s called big data or deep learning or AI" Um, I'm sorry, but unsupervised learning and deep learning are not the same.

Yeah, but garbage in is still garbage out. Which is the point he was trying to make.

It's not "garbage" it called science & mathematics, those terms have "meaning", and have lead to progress in hard long-standing problems, which have in turn lead to billions of dollars and millions of man-hours allocated to understanding and using them.

Just because you lack ability to understand nuances of something does not makes it "garbage".

Re: Deep-Fried Data

#73

I don't think I understand what the exact point of this talk was. Maybe the thesis was stated at the end of the talk when he said that he wishes the internet were more like a city rather than a mall. I think the internet can be like a city, and I think a great example of a place where people with conflicting ideas talk together is HN. Sure HN can be an echo chamber at times. But there's quite a few times when people…

The audience for this talk was a bunch of librarians and fellow travelers who are bringing large archives and collections online, often at great expense. I wanted to encourage them to find new, engaged audiences for these collections, rather than fixate on how to analyze them with computers. With regard to the dangers of surveillance, I've made a sustained argument about this in other talks. It boils down to the data…

Thank you for explaining that! The context is meaningful and makes your talk make sense.

On the regard of data talking into the wrong hands, I take issue to this argument because it's not a unique problem to personal data collection. Any data could be hacked - bank information, address, whatever. But that doesn't mean we don't use the internet for banking and etc. It means we try to make systems that are difficult to hack. It seems like you'd want data collection not to happen on websites like Facebook and Google, when hacking isn't a unique problem to those websites.

Re: Deep-Fried Data

#74

Earlier quoted context omitted.

Unsupervised refers to whether or not the dataset is being trained against anything. Think about the difference between: How many people will view this webpage? Divide these pages into 20 clusters? The first is supervised. The second isn't. Deep learning refers to a particular type of a particular learning technique: Specifically a neural network that has many hidden (intermediate) layers. Deep learning can be used f…

I agree with your sentiment, it feels out of place because deep learning, AI and big data are buzzwords, but unsupervised learning is a rather technical term in machine learning referring to a very specific class of problems.

Deep learning isn't a buzzword, it has a fairly precise definition. It describes a particular class of algorithms that happens to be the state of the art for many problems.

Re: Deep-Fried Data

#75

Earlier quoted context omitted.

The audience for this talk was a bunch of librarians and fellow travelers who are bringing large archives and collections online, often at great expense. I wanted to encourage them to find new, engaged audiences for these collections, rather than fixate on how to analyze them with computers. With regard to the dangers of surveillance, I've made a sustained argument about this in other talks. It boils down to the data…

Thank you for explaining that! The context is meaningful and makes your talk make sense. On the regard of data talking into the wrong hands, I take issue to this argument because it's not a unique problem to personal data collection. Any data could be hacked - bank information, address, whatever. But that doesn't mean we don't use the internet for banking and etc. It means we try to make systems that are difficult to…

Here's a capsule summary of what I'm pushing for: http://idlewords.com/six_fixes.htm

Re: Deep-Fried Data

#76
post #66
post #51

Earlier quoted context omitted.

> And I don't mean open source projects. Then what do you mean? You described exactly what some of the largest, most successful FOSS projects (Firefox, KDE, Gnome, Libre Office, FreeBSD) are already doing. > Let's group together smart people wanting to make a difference and have a hit list of things we (people) actually need. Well, the FSF maintains a list of "high priority Free Software projects" that need help, but…

> You described exactly what some of the largest, most successful FOSS projects (Firefox, KDE, Gnome, Libre Office, FreeBSD) are already doing. They "make a difference"? How exactly? At best, I can understand that for Firefox.

By providing quality software that lets me and thousands of others get useful work done, and not forcing us to accept onerous licensing terms in the process?

But hey, none of those can cure cancer, so what's the point, right?

Re: Deep-Fried Data

#77

Earlier quoted context omitted.

Thank you for explaining that! The context is meaningful and makes your talk make sense. On the regard of data talking into the wrong hands, I take issue to this argument because it's not a unique problem to personal data collection. Any data could be hacked - bank information, address, whatever. But that doesn't mean we don't use the internet for banking and etc. It means we try to make systems that are difficult to…

Here's a capsule summary of what I'm pushing for: http://idlewords.com/six_fixes.htm

I agree with that, but I have a small suggestion until these points are reality: you can consider adding/enabling SSL/TLS for your blog. Thanks!

P.S. I really like your posts and your tweets are hilarious, please don't ever stop.

Re: Deep-Fried Data

#78
post #27

Earlier quoted context omitted.

Yeah, but garbage in is still garbage out. Which is the point he was trying to make.

It's not "garbage" it called science & mathematics, those terms have "meaning", and have lead to progress in hard long-standing problems, which have in turn lead to billions of dollars and millions of man-hours allocated to understanding and using them. Just because you lack ability to understand nuances of something does not makes it "garbage".

Respectfully, I believe the "garbage" to which yarou was referring to was not the algorithms, but simply to the data that is being fed into these algorithms.

The point being that, no matter how sophisticated these techniques are, the quality of the results is constrained by the quality of the input data.

As Charles Babbage said: "On two occasions I have been asked, 'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?' I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question."

Re: Deep-Fried Data

#79

Machine learning does not have less bias than human researchers. It is simply magnified at scale. And that scale is exactly the state of the internet. There is so much data available to study and understand, that we absolutely need better tools, like machine learning or whatever we want to call it, to help us keep up. Shit's moving faster than our human perception can handle, especially for those who didn't grow up w…

> Machine learning does not have less bias than human researchers.

You are right that machine learning gains bias from the humans that created it, but unless they managed to transfer 100% of their biases to it, it will always have less bias.

Re: Deep-Fried Data

#80
post #63

Machine learning does not have less bias than human researchers. It is simply magnified at scale. And that scale is exactly the state of the internet. There is so much data available to study and understand, that we absolutely need better tools, like machine learning or whatever we want to call it, to help us keep up. Shit's moving faster than our human perception can handle, especially for those who didn't grow up w…

> Machine learning does not have less bias than human researchers. It is simply magnified at scale. Scale differences can and often do lead to qualitative differences. Individual (or aggregate) human researchers are not hooked up in huge services to make inferences and deductions automatically about billions of people. Besides those machine learning tools, beside the huge data sets, are programmed in their general fr…

Individual (or aggregate) human researchers are not hooked up in huge services to make inferences and deductions automatically about billions of people

Yes they (we) are. It's the same data set. TV, movies, papers, internet videos et al. is all the same biased, labeled data that is being fed (watched, listened to etc...) to machines. You automatically make inferences and deduce things about people based on labeling and training of your brain. You're constantly fine tuning by getting new weights about things through interactions with others and media.

Post reply on HN