workinengineering
BrowseField notes
workinengineering

Curated engineering jobs in Europe. Built to support the career path of engineering people at European companies across the EU, UK and EMEA.

Jobs

  • Browse all jobs
  • Remote jobs
  • Jobs in Germany
  • Jobs in Netherlands
  • Jobs in the UK
  • Field notes

About us

  • About
  • Contact
  • Privacy
  • Terms
  • Cookies

© 2026 workinengineering.eu — Made in Europe.

  • LinkedIn
  • X
  • Email
workinengineering
BrowseField notes
  1. All roles
  2. Stability AI
  3. Senior Site Reliability Engineer
Stability AI logo
Stability AI

Senior Site Reliability Engineer

RemoteSeniorRemote · 2 countriesApplications closedInfrastructureAI/ML

Our Take

Build and improve cloud infrastructure for Stability AI, a leader in generative AI.

What you’ll do

  • Develop and enforce SRE best practices
  • Architect and manage scalable cloud systems
  • Implement infrastructure as code (Terraform)
  • Set up monitoring, logging, and alerting
  • Drive incident management and root cause analysis
  • Mentor junior team members

What they’re looking for

  • Experience scaling resource-intensive systems
  • Knowledge of Kubernetes or container scaling
  • Background in software development/automation
  • Experience with Grafana, ELK stack, or similar
  • Cloud security experience
  • AWS and other cloud environments

Skills & Focus Areas

  • AWS
  • Terraform
  • Kubernetes
  • Grafana
  • ELK stack
  • CI/CD
  • Infrastructure as Code

As posted by Stability AI

< Remote - United States >

Job Description:
Stability AI’s Engineering Operations team is looking for a Senior Site Reliability Engineer (SRE) to join our growing team and play a pivotal role in improving and shaping our cloud infrastructure. The person will closely work with engineering, IT, security, and product teams to drive innovation and reliability in an evolving environment. Candidates should have the initiative to build and improve a maturing cloud landscape.

Responsibilities:

  • Developing and enforcing SRE best practices and standards across the organization.
  • Architecting and managing scalable systems in AWS and other cloud environments, focusing on high availability and resilience.
  • Implementing and maintaining infrastructure as code using Terraform.
  • Setting up and refining monitoring, logging, and alerting systems.
  • Driving incident management and root cause analysis to improve system reliability.
  • Championing SRE principles and mentoring junior team members.

Qualifications:

  • Collaborating with development teams to enhance CI/CD pipelines.
  • Experience scaling resource intensive systems, be it storage, networking, or compute.
  • Knowledge and experience with Kubernetes or other container scaling solutions
  • Background in software development or automation scripting.
  • Knowledge and experience with Grafana, ELK stack, or similar tools.
  • Cloud security experience.

Equal Employment Opportunity:

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.

 

Heads up: this role is no longer open for new applications.

We keep the page live for context and search continuity, but the apply action has been disabled.

Location
Remote: the United Kingdom (+1 non-EU)
Work mode
Remote
Track
IC (Individual Contributor)
Seniority
Senior
PostedApr 25
Stability AI logo
Stability AI
stability.ai
View all 2 Stability AI roles →

Was this listing helpful?