Live data from Hacker News

Deep-Fried Data

idlewords.com

11–20 of 149 posts

Re: Deep-Fried Data

#11
post #2

"...Dim witted grad student that you can't really trust..." Reminds me of the phrase "graduate student descent" for training neural networks... I've been noticing more casual dismissiveness towards grad students lately. They are certainly often treated as the grunt laborers of academia, in areas where career prospects are downright stupid. I generally feel it would be more productive to at least pretend that they're…

grad students are put in the same category as interns and teenagers, a naive type of person still in the making. i dont think there's any ill will intended.

Re: Deep-Fried Data

#12
If you think the Internet is as safe and controlled as a shopping mall, you probably should be reading Krebs on Security more.

People tend to move towards the more mall-like areas of the Internet due to spam and abuse that they don't want to deal with. This can be low-level stuff, or (as in the cases of Kreb himeself) sometimes the attackers get out the big guns, and you need to run for cover.

And that's why we're hanging out here, after all, and not in some unmoderated forum. And even here, post on certain subjects and conversation quickly degenerates.

I think we do need a wider variety of spaces to hang out, though. No set of rules works for everyone. And if you do want 4chan, you know where to find it.

Re: Deep-Fried Data

#13
I'm currently applying for co-op jobs (internships) and while trawling the university job board I've seen many positions requiring big data this or machine learning that.

What's not clear to me is why companies who don't seem to have any need for machine learning team (i.e. a subscription box company) are looking to hire one.

Surely part of this can be pinned down to the hype associated with ML that may well die out, but the proliferation of these tools doesn't bode well for Maciej's dream of a weird, creative, and interesting internet.

Re: Deep-Fried Data

#14
post #13

I'm currently applying for co-op jobs (internships) and while trawling the university job board I've seen many positions requiring big data this or machine learning that. What's not clear to me is why companies who don't seem to have any need for machine learning team (i.e. a subscription box company) are looking to hire one. Surely part of this can be pinned down to the hype associated with ML that may well die out,…

These companies aren't looking for someone to develop new machine learning techniques, they are just looking for someone who can slap together existing utilities to meet their goals.

Companies that run on subscription literally live and die by their churn rate. It is both feasible and reasonable for a subscription box company to hire someone to use machine learning to build a predictive churn model. That may seem trivial to you but that's the reality behind those job posts.

Re: Deep-Fried Data

#15
Regarding Maciej's fears about machine learning -

I've written about this before, and even right now I'm not sure where I stand exactly, except that tweaking the algorithms to compensate for bias is definitely not the right answer: if you look at the mirror and don't like what you see, you don't draw on top of the mirror to accentuate the result! You go on a diet!

I liked the idea of data gardening, but the thought of going-to-communities is daunting. I get tired even thinking about it.

Regarding living beyond walled gardens:

> Publish your texts as text. Let the images be images. Put them behind URLs and then commit to keeping them there. A URL should be a promise.

But people already do that! The question now is to turn to why people do otherwise. I personally do not understand the reason people say, post long blogposts on Facebook, but I do understand for services like Medium.

For example, I'm extremely tempted to write on Medium because it provides the network effects of readers clicking on tags to read next. So the question is how do we democratize that?

Re: Deep-Fried Data

#16
post #6

I provide guidance to attorneys involved in the discovery process; "Technology Assisted Review" is of huge interest to those teams, as it allows them to leverage coding on a small sample of the population across a much larger set of documents. For many cases, the cost and (occasional) time savings is instantly attractive. Sadly, the process is hard to do well. Far too many screw it up in new and amazing ways. The aut…

The issue with that approach is ensuring the suppressed are represented. When it's black vs white, you can oversample one and be done.

However, if there's any winner take all built into the system, there's a strong incentive to not even acknowledging dissent.

Re: Deep-Fried Data

#17
Have to admit that I didn't expect to see that quirk of LiveJournal culture mentioned in an article on the HN front page, let alone in a speech to the Library of Congress. It just sort of faded away without really influencing the current generation of social networks.

Also, it's funny how the net changes, how unthinkable it is to have a social network that doesn't slice up people's data and use it to advertise to them now compared to how anti-advertising LiveJournal was back then. Not convinced it's a change for the better.

Re: Deep-Fried Data

#18
post #2

"...Dim witted grad student that you can't really trust..." Reminds me of the phrase "graduate student descent" for training neural networks... I've been noticing more casual dismissiveness towards grad students lately. They are certainly often treated as the grunt laborers of academia, in areas where career prospects are downright stupid. I generally feel it would be more productive to at least pretend that they're…

Considering the working conditions and prospects for the future that graduate students face, one /could/ argue that that a selection bias should be expected there.

> to at least pretend that they're being trained to be independent, aggressive researchers

But that is the issue, isn't it-- it would be pretending.

Re: Deep-Fried Data

#19
> Many [programmers] work jobs that are intellectually stimulating, but ultimately leave nothing behind. There is a large population of technical people who would enjoy contributing to something lasting.

This hits pretty close to home.

Re: Deep-Fried Data

#20

Earlier quoted context omitted.

What's the distinction?

Unsupervised refers to whether or not the dataset is being trained against anything. Think about the difference between: How many people will view this webpage? Divide these pages into 20 clusters? The first is supervised. The second isn't. Deep learning refers to a particular type of a particular learning technique: Specifically a neural network that has many hidden (intermediate) layers. Deep learning can be used f…

I agree with your sentiment, it feels out of place because deep learning, AI and big data are buzzwords, but unsupervised learning is a rather technical term in machine learning referring to a very specific class of problems.
Post reply on HN