Live data from Hacker News

Deep-Fried Data

idlewords.com

61–70 of 149 posts

Re: Deep-Fried Data

#61
I don't think I understand what the exact point of this talk was. Maybe the thesis was stated at the end of the talk when he said that he wishes the internet were more like a city rather than a mall. I think the internet can be like a city, and I think a great example of a place where people with conflicting ideas talk together is HN. Sure HN can be an echo chamber at times. But there's quite a few times when people with differing opinions talk about their different opinions.

Also I don't necessarily understand Ceglowski's stance on why we shouldn't use deep learning and should avoid surveillance on the web. I don't take issue with becoming a datapoint in Facebook's web of people because nothing bad has happened or can happen from me giving Facebook my data. When most people speak out about the data that's being collected about Facebook and Google users they say they're "worried about what could happen" but then never list any bad things that they're actually afraid of. The speaker falls to this issue too. Ceglowski says:

>I worry about legitimizing a culture of universal surveillance.

But then never explains what bad could happen from legitimizing that culture. Maybe I'm completely missing the point of the talk? Please explain what I'm missing if I'm actually missing something.

Re: Deep-Fried Data

#62

Earlier quoted context omitted.

My degree was in studio art, not art history.

Those two are still orders of magnitude closer to each other when compared to difference between unsupervised learning with deep learning.

As someone with not particularly deep knowledge in either area, that admittedly sounds a lot like "no, but the differences between subfields in MY subfield are way more important than all that stuff over there", which is similar to what you claim the article does. They are important once you care about any details, but not for just describing changing fashions.

So I'm curious to hear a good explanation for that assertion, founded in knowledge of both areas.

Re: Deep-Fried Data

#63

Machine learning does not have less bias than human researchers. It is simply magnified at scale. And that scale is exactly the state of the internet. There is so much data available to study and understand, that we absolutely need better tools, like machine learning or whatever we want to call it, to help us keep up. Shit's moving faster than our human perception can handle, especially for those who didn't grow up w…

>Machine learning does not have less bias than human researchers. It is simply magnified at scale.

Scale differences can and often do lead to qualitative differences.

Individual (or aggregate) human researchers are not hooked up in huge services to make inferences and deductions automatically about billions of people.

Besides those machine learning tools, beside the huge data sets, are programmed in their general framework by human researchers, and are given weights, constraints, and fine-tuning by them, so they have both kinds of biases.

>But sure demonizing the things you don't like is one step on the path to learning what's truly valuable.

So, kind of like disparaging via a straw-man a speech that offers detailed argumentation?

Re: Deep-Fried Data

#64
post #58

> I’ve saluted the efforts of Archive Team and the Internet Archive, but their activity is like having a museum curator that rides around in a fire truck, looking for burning buildings to pull antiques from. It's heroic, it's admirable, but it’s no way to run a culture. ...but in the meantime, here's an obligatory and shameless plug for donating to the Internet Archive[1] (tax-deductible in the US), or better yet mak…

The Archive is awesome, but the author's sensationalist description of what they do isn't really accurate. For the most part, archive.org is not rushing in to save stuff that's about to be deleted. Instead they are crawling the web 24/7, patiently maintaining a historical record. Check out http://oldweb.today It is amazing

[deleted]

Re: Deep-Fried Data

#65
post #58

> I’ve saluted the efforts of Archive Team and the Internet Archive, but their activity is like having a museum curator that rides around in a fire truck, looking for burning buildings to pull antiques from. It's heroic, it's admirable, but it’s no way to run a culture. ...but in the meantime, here's an obligatory and shameless plug for donating to the Internet Archive[1] (tax-deductible in the US), or better yet mak…

The Archive is awesome, but the author's sensationalist description of what they do isn't really accurate. For the most part, archive.org is not rushing in to save stuff that's about to be deleted. Instead they are crawling the web 24/7, patiently maintaining a historical record. Check out http://oldweb.today It is amazing

> For the most part, archive.org is not rushing in to save stuff that's about to be deleted.

Parent mentioned both the Internet Archive and Archive Team. You're right about the Internet Archive, but "rushing in to save stuff that's about to be deleted" is a pretty apt description of most of Archive Team's activity.

Re: Deep-Fried Data

#66
post #51
post #44

Earlier quoted context omitted.

Same here. I'm battling with this thought a lot. Beyond jobs, I think there should be communities of developers, designers, producers, writers, getting together and figuring out this stuff. And I don't mean open source projects. Let's group together smart people wanting to make a difference and have a hit list of things we (people) actually need. A group that would organise people into mission driven development. I'm…

> And I don't mean open source projects. Then what do you mean? You described exactly what some of the largest, most successful FOSS projects (Firefox, KDE, Gnome, Libre Office, FreeBSD) are already doing. > Let's group together smart people wanting to make a difference and have a hit list of things we (people) actually need. Well, the FSF maintains a list of "high priority Free Software projects" that need help, but…

>You described exactly what some of the largest, most successful FOSS projects (Firefox, KDE, Gnome, Libre Office, FreeBSD) are already doing.

They "make a difference"? How exactly? At best, I can understand that for Firefox.

Re: Deep-Fried Data

#67
post #58

> I’ve saluted the efforts of Archive Team and the Internet Archive, but their activity is like having a museum curator that rides around in a fire truck, looking for burning buildings to pull antiques from. It's heroic, it's admirable, but it’s no way to run a culture. ...but in the meantime, here's an obligatory and shameless plug for donating to the Internet Archive[1] (tax-deductible in the US), or better yet mak…

The Archive is awesome, but the author's sensationalist description of what they do isn't really accurate. For the most part, archive.org is not rushing in to save stuff that's about to be deleted. Instead they are crawling the web 24/7, patiently maintaining a historical record. Check out http://oldweb.today It is amazing

What the Archive Team saves does get uploaded to the Internet Archive, but they aren't officially part of it. I think Maciej's description of the Archive Team is accurate - they are the archivists of last resort. When a commercial service is about to disappear forever, they're the ones that spring into action and rescue as much data as possible. If companies and the people that comprise them cared enough about their users' data, there would be no need for the Archive Team.

Re: Deep-Fried Data

#68

I don't think I understand what the exact point of this talk was. Maybe the thesis was stated at the end of the talk when he said that he wishes the internet were more like a city rather than a mall. I think the internet can be like a city, and I think a great example of a place where people with conflicting ideas talk together is HN. Sure HN can be an echo chamber at times. But there's quite a few times when people…

The audience for this talk was a bunch of librarians and fellow travelers who are bringing large archives and collections online, often at great expense. I wanted to encourage them to find new, engaged audiences for these collections, rather than fixate on how to analyze them with computers.

With regard to the dangers of surveillance, I've made a sustained argument about this in other talks. It boils down to the data being collected having great power to harm people if it is ever put to malicious use, and a lifespan that exceeds that of institutions we know how to run. My beef is not with the surveillance alone, but with the combination of surveillance and permanent storage.

Re: Deep-Fried Data

#69

> Publish your texts as text. Let the images be images. Put them behind URLs and then commit to keeping them there. I sounds like he's saying ephemeral content is worthless and should be shunned. I, and hundreds of millions of others, disagree. You want a bland, awful, boring society? Easy: make everything you do stick around forever—like a promise. And then watch the world self-police as the lifeblood drains out of…

The audience for this talk was people with very large collections they're bringing online. I was trying to encourage them to avoid exotic formats, custom plugins, custom software (shudder) when they put this material online, and make them web accessible.

For example, here is three quarters of a PETABYTE of historical American newspapers: http://chroniclingamerica.loc.gov

Re: Deep-Fried Data

#70
post #54

Earlier quoted context omitted.

The difference doesn't really matter in context. You're fixating on a small part of the article that isn't important to the main thread.

Its not a "small part", its a basic litmus test. The four terms are completely different from each other, and are not names of methods. Unsupervised learning: Learning without a set of labels. Big Data: Collecting / using large amount of data. Deep Learning: Complex, multilayer representations which perform better than shallow/linear representations. AI: Artificial Intelligence, an overarching subject or grouping of…

It's important to understand that this is not a technical talk/article and providing those examples in a sense "there are data, people analyze it, here are some stuff you might've heard" is fine.

You wouldn't complain that someone mentioned astronomy and music as an example in the talk about education, even though those are quite different disciplines.

Post reply on HN