Live data from Hacker News

"SRE" doesn't seem to mean anything useful any more

rachelbythebay.com

31–40 of 84 posts

Re: "SRE" doesn't seem to mean anything useful any more

#31
post #10

Earlier quoted context omitted.

DevOps has that meaning to you, maybe. I would implore some reading into the origin. Cliff notes version, things that happened around the same time: * Dev+Ops days is made by Patrick Dubois, the intent was for Sysadmins to use Agile ("Agile Systems Administrator" was the desired job title), he used "dev"+"ops" to signify a unification in working methodology, not as a team * 10+ deploys a day from Flickr; where the au…

You need to also look into why DevOps became a thing. Because devs were sick of the BS dealing with Ops teams (want a server, that will take 6 months to provision). So devs decided they could do it themselves (with better tools).

Perhaps they would’ve spent more effort making software not suck if they still had that constraint. Instead, now we get the modern dumpster fire because more resources are an HPA call away.

Re: "SRE" doesn't seem to mean anything useful any more

#32
post #2

People need operations staff, people don't like operations staff and keep trying to treat them like developers. But, operations staff do and have always developed software, just internal software for glue or orchestration, and they work differently to regular software developers in that their customers are usually themselves to meet an internal objective of reliability, stability or ease-of-use for developers. It's i…

Same here, being that developer on the team that would embrace build and deployment scripts, means that for the last 30 years, I have always to jungle around, am I a developer, a systems administrator, network engineer, or whatever is fashionable as term for upper management.

The worse part of the fashionable job titles is that both management and HR love to put people on little boxes, and generalists are a big headache for them, on top of deteriorating the actual original meanings of these roles.

Re: "SRE" doesn't seem to mean anything useful any more

#33

Earlier quoted context omitted.

> So devs decided they could do it themselves (with better tools). And failed. Other people’s jobs always look easy from the outside

Actually devs focusing on Ops lead to development of the automation we use today for most Ops activities, so I would say that is a success.

We already had such automation tools under system administration umbrella.

What do we think we were using Perl for?

Re: "SRE" doesn't seem to mean anything useful any more

#34
post #32
post #2

People need operations staff, people don't like operations staff and keep trying to treat them like developers. But, operations staff do and have always developed software, just internal software for glue or orchestration, and they work differently to regular software developers in that their customers are usually themselves to meet an internal objective of reliability, stability or ease-of-use for developers. It's i…

Same here, being that developer on the team that would embrace build and deployment scripts, means that for the last 30 years, I have always to jungle around, am I a developer, a systems administrator, network engineer, or whatever is fashionable as term for upper management. The worse part of the fashionable job titles is that both management and HR love to put people on little boxes, and generalists are a big heada…

I used to work at a startup during the 2k tech bubble and a coworker had the title "General Specialist" on their business cards.

It's something I still aspire to be :)

Re: "SRE" doesn't seem to mean anything useful any more

#35

Honest question, what would you call a role that: * Is on call * Manages internal software (grafana, Prometheus, salt stack, etc) * is the first line of defense for issues in the field, works with support and the engineering team to handle problems * Manages a distributed fleet of servers (uses off the self and/or custom code to do so) * Builds internal tools/automations to improve the reliability of our platform and…

"Platform Engineer" has started to grow as a name for pretty much this. An engineer that builds the platform that all things run on. For example, the k8s clusters, the observability stack, has on-call, and builds the internal tools and automations (called the IDP, or internal developer platform).

IDP already has a very important meaning as a part of any platform :)

Re: "SRE" doesn't seem to mean anything useful any more

#36

Honest question, what would you call a role that: * Is on call * Manages internal software (grafana, Prometheus, salt stack, etc) * is the first line of defense for issues in the field, works with support and the engineering team to handle problems * Manages a distributed fleet of servers (uses off the self and/or custom code to do so) * Builds internal tools/automations to improve the reliability of our platform and…

What I've experienced working well is to consider platform/infra as just another dev team. They should experience the same good practices (staging changes, tests, documentation, clean code) and duties (level 2 on calls, regular postmortems from ops, etc).

Ops/prod/support eng are generally more business related and less technical than infra/platform. They make sure the processes run as they should, or operate with agreed procedures when something is raised. Some issues may be lightly technical ("theres an alert on free disk space of server X", "provider Y changed their SFTP keys without notice", etc) and some business related ("provider Z is late to push updates", "it's a holiday on country of provider A", "service of team B raises an exception when situation C occurs"). They are often level 1, in between the actual dev teams (escalating issues to them, or asking for better stability, logs, resiliency, etc) and the platform/infra team.

Being faced with production issues is a burden for focus, etc. I wouldn't want my platform team to be on level 1 handling these stuff or nothing get done. Still, I wouldn't want them to be completely free from it or they would loose track with reality, just as any other dev team.

To summarize, I would say:

- A fleet of business dev team - A fleet of platform / infra team - A support/ops/prod team handling level 1 through procedures written by dev/platform teams, and raising to infra or business dev for level 2. Regular feedback sessions with each other team to have them stay on touch with real life. Some light coding / improvements tasks paired with infra or dev to get them a better understanding of the underlying layers.

From the description you provided, that's what I would call a "platform engineer". Though the titles are sensitive for some people, and there's a lot of discrepancies between companies, so i usually let people from platform choose the title they prefer, between SRE / DevOps / platform eng / cloud eng / infra eng.

Re: "SRE" doesn't seem to mean anything useful any more

#37

Honest question, what would you call a role that: * Is on call * Manages internal software (grafana, Prometheus, salt stack, etc) * is the first line of defense for issues in the field, works with support and the engineering team to handle problems * Manages a distributed fleet of servers (uses off the self and/or custom code to do so) * Builds internal tools/automations to improve the reliability of our platform and…

[deleted]

Re: "SRE" doesn't seem to mean anything useful any more

#38

Honest question, what would you call a role that: * Is on call * Manages internal software (grafana, Prometheus, salt stack, etc) * is the first line of defense for issues in the field, works with support and the engineering team to handle problems * Manages a distributed fleet of servers (uses off the self and/or custom code to do so) * Builds internal tools/automations to improve the reliability of our platform and…

It used to be systems administrator, or network engineer, depending how deep into infrastructure it goes.

Re: "SRE" doesn't seem to mean anything useful any more

#39

Honest question, what would you call a role that: * Is on call * Manages internal software (grafana, Prometheus, salt stack, etc) * is the first line of defense for issues in the field, works with support and the engineering team to handle problems * Manages a distributed fleet of servers (uses off the self and/or custom code to do so) * Builds internal tools/automations to improve the reliability of our platform and…

As long as that very description is spelled out clearly near the top of the job ad, it doesn't matter terribly much (within reason) what you call it - job searchers will try various different strings to find it. Personally, I'd call it an ops role, however. In general, most of the issues I've seen with these sorts of roles aren't the naming of them but rather the third bullet in your list. Having an ops role that's r…

I haven’t started trying to hire for this position yet, it’s been something I’ve been thinking of for a while though. Right a small number of developers, myself included, are on call and I’d like to reduce that a bit or just share the burden.

My goal isn’t to throw crap over the fence and say “make it work” but rather empower someone to make maintaining and growing our platform their main goal. The developers (again, myself included) are not great at the ops side of things and can rarely focus on the infrastructure itself due to other priorities (yes, we can talk about how that itself is an issue). If I could clone myself and one of specialize in ops and the other on programming for the platform I would in a heartbeat.

Infra/Ops and programming are two different mindsets (much like managing people or qa differs from writing code). Switching between them is hard and you pay a penalty to do so. Not to mention there are skills (networking is high on that list) that I’m not good at. I can scrape by but that’s not where my skills lie. That’s why I’d like to hire someone who is good at it, who _does_ enjoy it, and who push for changes from a ops perspective that I can’t due to time or skill.

Re: "SRE" doesn't seem to mean anything useful any more

#40
It doesn’t mean anything for two reasons: companies have treated it as a catch—all, and there is a glut of people calling themselves SREs who have never operated a server that wasn’t in a cloud.

You can learn enough about Linux to be decent at your job on only VMs if you’re dedicated, but I’d argue that until you’ve also dealt with hypervisors, bare metal, and hardware issues, you’re missing some of the picture.

“That’s no longer applicable, so why should I care?” Because it pops up everywhere I’ve been. Random build server that everyone forgot about but is critical suddenly shits itself, it’s running some ancient version of Ubuntu, and it’s all hand-rolled. Someone decided to provision a bunch of EC2s with bash, but they made critical errors like not knowing to make a new initramfs after configuring mdadm, so now the RAID disappears on reboots.

Understanding the fundamentals has always and will always matter. Anyone telling you differently is selling you something.

Post reply on HN