
Backend Engineer II - Data Platform
Our Take
Backend Engineer II at Spotify building and operating core backend services and APIs for the data platform.
What you’ll do
- Design, build and operate the backend services and APIs behind our data catalog, lineage and publishing systems
- Evolve backend services and APIs as the platform adopts new storage technologies and new kinds of data
- Develop and run the batch and streaming pipelines that keep our metadata current
- Process platform events and audit logs to capture how data is created and changed
- Turn lineage and metadata into data products other teams build on
- Turn what the teams using our platform need into technical designs
- Drive technical designs through delivery and rollout
- Own the results once services are in production
What they’re looking for
- Solid experience building and running backend services and APIs in production with Java
- Comfortable with SQL
- Experience with data pipelines or a cloud data warehouse such as BigQuery, Snowflake or Databricks
- Working understanding of data engineering concepts such as orchestration, metadata management, data quality, governance and lineage
- Strong communicator and collaborator
What you get
- Flexibility to work where you work best
- In person meetings are required
Skills & Focus Areas
- Java
- SQL
- BigQuery
- Snowflake
- Databricks
- Beam
- Flink
- Spark
As posted by Spotify
Every dataset at Spotify, from the data behind creator royalties to the signals powering recommendations, is registered, published and traced through systems our team builds. We're the source of truth for what data exists at Spotify, where it lives, who owns it, and how it flows from one pipeline to the next.
These systems sit on the critical path of the platform. Data teams use them every day to publish and find data, and other platform teams build on them for access, retention, governance and incident response. That means our work is judged on correctness and reliability, because every system built on top of ours is only as good as the metadata we provide.
This is a backend role focused on data management, working on the problems at the core of any large data platform: running a data catalog that stays accurate at scale, capturing lineage across batch and streaming workloads, managing schemas as they evolve, and guaranteeing that consumers only read complete data. We build on open standards like OpenLineage and Apache Iceberg, and define the internal standards for naming, storage and metadata that the rest of Spotify builds against. As the platform takes on new storage technologies and new kinds of data assets, those standards and the systems behind them have to evolve with it, and you'll help decide how.