workinengineering
BrowseField notes
workinengineering

Curated engineering jobs in Europe. Built to support the career path of engineering people at European companies across the EU, UK and EMEA.

Jobs

  • Browse all jobs
  • Remote jobs
  • Jobs in Germany
  • Jobs in Netherlands
  • Jobs in the UK
  • Field notes

About us

  • About
  • Contact
  • Privacy
  • Terms
  • Cookies

© 2026 workinengineering.eu — Made in Europe.

  • Email
workinengineering
BrowseField notes
  1. All roles
  2. Encord
  3. Senior SRE Engineer
Encord logo
Encord

Senior SRE Engineer

OnsiteSeniorLondon, England, the United KingdomAI/MLPlatformInfrastructure
Apply directly at Encord

Our Take

Build and operate Encord's AI data platform infrastructure for 1000s of AI teams.

What you’ll do

  • Optimize services for large-scale data pipelines
  • Improve CI/CD pipelines
  • Define and own SLIs/SLOs/SLAs
  • Design, deploy, maintain cloud infrastructure
  • Build and manage Kubernetes clusters
  • Instrument services for observability
  • Guide automation and tooling efforts

What they’re looking for

  • Hands-on SRE/DevOps experience
  • Design/build resilient distributed systems
  • Networking, OS, database fundamentals
  • Observability: metrics, logs, traces
  • Comfortable with on-call/incident management
  • Experience with Kubernetes
  • Experience with GCP or AWS

What you get

  • Competitive salary and equity
  • 25 days annual leave + UK holidays
  • Annual learning & development budget
  • Travel for customer visits and events
  • Company lunches twice a week
  • Monthly socials & bi-annual offsites

Skills & Focus Areas

  • Kubernetes
  • GCP
  • AWS
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Datadog
  • CI/CD

As posted by Encord

About us

Encord is the universal data layer for AI that helps 300+ AI teams train and run models on the right data. Our platform indexes, curates, annotates, and evaluates data across the full AI lifecycle, from development through production.

 

Trusted by Woven by Toyota, AXA, UiPath, Zipline, and more. We're an ambitious team of 100+ working at the frontier of AI and have raised $60M in Series C funding from Wellington Management, CRV, Next47 and Y Combinator.

The role

We're looking for a Senior Site Reliability Engineer to join our growing platform engineering team. You'll be embedded in the teams building and operating Encord's core infrastructure, ensuring our platform is performant, reliable, observable, and scalable.

You will lead the planning and execution of efforts needed as we grow from our customer base from hundreds to thousands of AI teams worldwide, and the volume of AI training and supervision data managed by our platform from TBs to PBs of data.

You'll drive a culture of performant and resilient software through individual contributions and collaboration with multiple squads.

  

What You'll Do

  • Performance & Capacity — Profile and optimise services handling large-scale data pipelines; perform capacity planning for storage and compute-intensive workloads. Work with squads to establish performance benchmarks and expectations

  • Collaboration — Partner closely with backend and ML engineers to improve deployment pipelines (CI/CD), review infrastructure changes, and champion reliability best practices.

  • Reliability & Availability — Define and own SLIs/SLOs/SLAs for critical services; build alerting, runbooks, and incident response processes; lead postmortems with a blameless culture.

  • Infrastructure & Cloud — Design, deploy, and maintain cloud infrastructure on GCP and AWS; manage Kubernetes clusters, networking, and storage at petabyte scale.

  • Automation & Tooling — Work to improve developer productivity and guide and review automation and tooling efforts across the engineering group.

  • Observability — Instrument services with distributed tracing, logging, and metrics (Prometheus, Grafana, OpenTelemetry, Datadog or similar); build infrastructure, define best practices and work with each squad to ensure every service is observable before it goes to production.

  

What We're Looking For

  • Experience on hands-on SRE, DevOps, platform engineering experience or similar in a production environment.

  • Strong fundamentals in designing, building and maintaining resilient distributed and/or high performance systems

  • Solid understanding of networking, operating systems and database technologies

  • Experience with observability fundamentals — metrics, logs, traces, and alerting.

  • Comfortable with on-call rotations and incident management.

Tech stack

  • We are technology agnostic at Encord and not looking for experience across all of these — as long as you're open to learning, please apply.

    • Backend: Python and Rust

    • Frontend: TypeScript and React

    • Deployment: Kubernetes

    • Infrastructure: GCP

       

Why Encord

  • Competitive salary, commission, and meaningful equity in a high-growth startup

  • Strong in-person culture — most of the team works from our London office 4+ days/week

  • 25 days annual leave + UK public holidays

  • Annual learning & development budget

  • Travel for customer visits, events, and conferences across the UK and Europe

  • Company lunches twice a week

  • Monthly socials & bi-annual team offsites

Apply directly at Encord
Location
London, England, the United Kingdom
Work mode
Onsite
Track
IC (Individual Contributor)
Seniority
Senior
Type
FULL_TIME
PostedJun 29
Encord logo
Encord
encord.com
View all 9 Encord roles →

Was this listing helpful?