Live data from Hacker News

Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

engineering.atspotify.com

111–120 of 135 posts

Re: Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

#111
Regarding the title, how do we know this is THE largest dataflow job? The article body itself only makes mention that this is THEIR largest dataflow job. This post doesn't make any quantifiable claims either one could use to support, this is all I found:

> "We estimate around a 50% decrease in Dataflow costs this year compared to previous years’ Bigtable-based approach. Additionally, we avoided scaling the Bigtable cluster up two to three times its normal capacity (up to around 1,500 nodes at peak"

The official Spotify Engineering Tweet similarly only makes mention that this is Spotify's largest dataflow job ever: https://twitter.com/SpotifyEng/status/1359887825047613442.

I'm fairly sure a similar accidental unsourced exaggeration was made last year.

Maybe the title should be Spotify Optimized Their Largest Dataflow Job Ever For Wrapped 2020?

Re: Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

#114
post #61

Earlier quoted context omitted.

What's scary about this? They're complying with GDPR. Isn't that a good thing? The "scary" thing in that tweet is that they store the manufacturer of their bluetooth headphones?

That's not what one could call complying. He had to follow them like a dog for a long while just to get his rightful data.

That was from 2018. Almost 3 years ago.

Re: Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

#116

I'd be curious how this compares in load to Google's internal applications. I'm also curious what the capacity of Google's infrastructure goes to Google vs. GCE - has combined GCE usage even passed the compute needs of Google internally yet?

The only even remotely concrete information in this post is their input was 1PB, and they typically have 500 bigtable tablet servers. In 2008, Google said they processed 20PB per day through mapreduce jobs. For the last ten years the only thing they've said about the size of their public web index is that it is over 100PB.

Re: Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

#117
post #32

Earlier quoted context omitted.

I wonder if it'll meet the same fate as Netflix at some point, with publishing houses going for their own streaming services instead. I don't think so because the way that people consume music and the way they consume films and television are very different. With a film you might block out a few hours to watch that specific provider. With music you're more likely to want to interleave content from several providers a…

Speaking of interleaving, I wish Spotify understood how to interact with Concept Albums. Changing to the middle of The Wall for one song is jarring and I usually skip it. And then I worry skipping it is going to train the algorithm that I don't like Pink Floyd.

I never understood why for albums like this they even bother splitting up the tracks.

Re: Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

#118
post #26

Wrapped works so well because it panders to you. Everyone likes to be acknowledged for listening to the weird indie band they discovered earlier this year. I enjoy it anyways, and Spotify is still a great service for now - I wonder if it'll meet the same fate as Netflix at some point, with publishing houses going for their own streaming services instead.

If anything the top-5 format doesn't capture the long tail you might be listening to.

Re: Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

#119
This article fails to make a clean problem statement for the general audience. It jumps right into jargon and names from some framework/library. It reads like it was an internal report from a programmer to their team, and someone decided to make it public with no changes.

Re: Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020

#120

Earlier quoted context omitted.

> I'm convinced that it is now important to hold on to older appliances that work without internet access or data collection Or run modern, up to date FOSS equivalents on machines you control. I’ve migrated more and more services like that and I’m slowly but surely building my own “cloud”.

I understand why people don't constantly think about their digital footprint, but it's not that different from when they order the same thing at a stall for a month, and the clerk begins to ask "Do you want your usual?".

But the digital footprint is infinitely copyable, permanent etc
Post reply on HN