Managing Machines at Spotify
labs.spotify.com
Managing Machines at Spotify
1–10 of 20 posts
Re: Managing Machines at Spotify
#2Let me fix that for you; stop gendering your servers.
Re: Managing Machines at Spotify
#3"We also assigned each server a static unique identifier in the form of a woman’s name – a shrinking namespace with thousands of servers." Let me fix that for you; stop gendering your servers.
Re: Managing Machines at Spotify
#4It's strange to me that this is still so common. My theory is that the "one machine one port" philosophy is still built into a lot of software (monitoring, the ELB, etc). Another is that this is the philosophy we've always known.
Take a look at Kubernetes. Everything is accessible via localhost:. that breaks most home-built and enterprise orchestration and monitoring tools spectacularly even though it's a much simpler mode (everything is a port, not ip port combo).
Density is much easier to accomplish on larger machines with more cores, which are elastic in the face of bursty residents. They are also generally cheaper per compute/memory.
Re: Managing Machines at Spotify
#5Sorry for so many questions, but you made a big deal about how manual "DNS curation" was a bad thing, then glossed over the solution.
Re: Managing Machines at Spotify
#6Re: Managing Machines at Spotify
#7> While we heavily utilise Helios for container-based continuous integration and deployment (CI/CD) each machine typically has a single role – i.e. most machines run a single instance of a microservice. It's strange to me that this is still so common. My theory is that the "one machine one port" philosophy is still built into a lot of software (monitoring, the ELB, etc). Another is that this is the philosophy we've a…
IPv6 is practically built for containers, and, to Kubernetes's credit, they architected with that in mind. (Learned from BNS.) Weirdly, what I'm saying here was the original idea behind ports in the first place. There just aren't enough of them, particularly when half your space is shared with client sockets.
I want a world where v4 is pretty much just my control plane into the v6 cluster, since I'll die before IPv4. Google and far more importantly Amazon need to come up with a v6 story in their cloud offerings already. AWS has had a decade. This isn't just blind advocacy any more; the orchestration and software side is starting to build entire parts of the OSI stack because the network side of our industry is stuck without any sign of moving, no matter how dire the v4 situation.
Re: Managing Machines at Spotify
#8What happened to your DNS data? Did you switch to dynamic DNS based upon database data? You talk about how much a burden the manual DNS information was, but then you don't specify how you actually solved it using "automation." Is it all dynamic? Everything use SRV records that have TTL's or are added and removed? Sorry for so many questions, but you made a big deal about how manual "DNS curation" was a bad thing, the…
Re: Managing Machines at Spotify
#9> While we heavily utilise Helios for container-based continuous integration and deployment (CI/CD) each machine typically has a single role – i.e. most machines run a single instance of a microservice. It's strange to me that this is still so common. My theory is that the "one machine one port" philosophy is still built into a lot of software (monitoring, the ELB, etc). Another is that this is the philosophy we've a…
It's also harder to separate two or more processes that 'grew up' together in the same container/machine/vm.
Re: Managing Machines at Spotify
#10What happened to your DNS data? Did you switch to dynamic DNS based upon database data? You talk about how much a burden the manual DNS information was, but then you don't specify how you actually solved it using "automation." Is it all dynamic? Everything use SRV records that have TTL's or are added and removed? Sorry for so many questions, but you made a big deal about how manual "DNS curation" was a bad thing, the…
Look at the section "DNS Pushes".
Ideally we hope to provide some followup posts that go deeper into technical detail about key pieces of the stack (DNS, initramfs framework, job broker, GCP usage, etc).