workinengineering
BrowseField notes
workinengineering

Curated engineering jobs in Europe. Built to support the career path of engineering people at European companies across the EU, UK and EMEA.

Jobs

  • Browse all jobs
  • Remote jobs
  • Jobs in Germany
  • Jobs in Netherlands
  • Jobs in the UK
  • Field notes

About us

  • About
  • Contact
  • Privacy
  • Terms
  • Cookies

© 2026 workinengineering.eu — Made in Europe.

  • Email
workinengineering
BrowseField notes
  1. All roles
  2. Deutsche Telekom AG
  3. Industrial AI Cloud - Platform Engineer (REF5501C)
Deutsche Telekom AG logo
Deutsche Telekom AG

Industrial AI Cloud - Platform Engineer (REF5501C)

HybridBudapest, Budapest, HungaryAI/MLInfrastructure
Apply directly at Deutsche Telekom AG

Our Take

Company Description As Hungary’s most attractive employer in 2025 (according to Randstad’s representative survey), Deutsche Telekom IT Solut

What you’ll do

  • Build, configure, and maintain Kubernetes clusters
  • Design and operate NVIDIA AI software stack
  • Manage Helm charts, GitOps, and automation scripts
  • Troubleshoot, tune, and scale Kubernetes workloads
  • Develop and maintain CI/CD pipelines
  • Operate and enhance monitoring stacks
  • Build, optimize, and secure container images
  • Support distributed AI workloads

What they’re looking for

  • Kubernetes Certified Administrator (CKA) or equivalent
  • Strong CI/CD (Jenkins, GitLab) and GitOps experience
  • Proficiency with Helm charts and Kubernetes
  • Scripting in Python/Bash; IaC with Terraform/Ansible
  • Experience with Docker, Podman, and image scanning
  • Knowledge of Prometheus and Grafana monitoring
  • Experience running AI/HPC workloads at scale

What you get

  • Work on Europe's first industrial AI cloud
  • Cutting-edge technologies
  • Direct collaboration with NVIDIA experts
  • Training opportunities
  • Career progression
  • Hybrid working model

Skills & Focus Areas

  • Kubernetes
  • CI/CD
  • GitOps
  • Terraform
  • Ansible
  • Docker
  • Prometheus
  • Grafana

As posted by Deutsche Telekom AG

Company Description

As Hungary’s most attractive employer in 2025 (according to Randstad’s representative survey), Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group. The company provides a wide portfolio of IT and telecommunications services with more than 5300 employees. We have hundreds of large customers, corporations in Germany and in other European countries.

DT-ITS recieved the Best in Educational Cooperation award from HIPA in 2019, acknowledged as the the Most Ethical Multinational Company in 2019. The company continuously develops its four sites in Budapest, Debrecen, Pécs and Szeged and is looking for skilled IT professionals to join its team.

Job Description

General description/ Purpose

NVIDIA and Deutsche Telekom are jointly developing the world’s first industrial AI cloud for European manufacturers. This AI factory in Germany will host 10,000 GPUs across NVIDIA DGX B200 systems and RTX Pro Servers. Deutsche Telekom provides secure, sovereign and fast infrastructure, including data centers, operations, security, and AI solutions.

Role Overview

We are seeking a Platform Engineer to build, automate, and operate the platform services of the Industrial AI Cloud. This role focuses on running and evolving large-scale Kubernetes clusters, container orchestration platforms, CI/CD pipelines, GitOps and associated automation to support AI workloads. Experience with Infrastructure as Code. You’ll be part of the team supporting Kubernetes workloads, ensuring smooth operations and continuous improvement of the platform layer, while collaborating with infrastructure, security, and AI teams.

Key Responsibilities

  • Operate & Evolve Kubernetes Platform: Build, configure, and maintain bare metal hosts and Kubernetes clusters to run GPU/AI workloads.
  • Design & Operate NVIDIA AI related software stack (Slurm, Run AI)
  • Provide customized application support for AI related workloads
  • Container Orchestration & Automation: Manage Helm charts, GitOps workflows, Ansible scripts, possibly Terraform code and automation for deploying services and AI workloads.
  • Operate Kubernetes Workloads: Act as primary contact for all Kubernetes-related topics, including troubleshooting, performance tuning, and scaling.
  • CI/CD & GitOps: Develop and maintain CI/CD pipelines with Jenkins and GitLab; implement GitOps practices for consistent deployments and infrastructure changes. Terraform basics.
  • Monitoring & Observability: Operate and enhance Prometheus and Grafana monitoring stacks for bare metal hosts, Kubernetes and platform services.
  • Container Images & Registries: Build, optimize and secure container images (Docker, Podman); manage registries and versioning, image scanning (Trivy).
  • Object Storage & Persistent Volumes: Integrate and maintain object storage solutions for AI workloads.
  • Run AI & HPC Workloads: Support and operate distributed AI workloads within bare metal hosts and Kubernetes environments.
  • Collaboration with Infrastructure & AI Teams: Coordinate closely with Infrastructure Engineers, data center staff and AI developers to ensure smooth delivery of services.
  • ITIL Processes: Follow incident, problem, and change management workflows; create and maintain operational runbooks. Adhere to ZERO outage guidelines.

What We Offer

  • Work on Europe’s first industrial AI cloud with cutting-edge technologies.
  • Direct collaboration with NVIDIA and Deutsche Telekom experts.
  • Hybrid working model, training opportunities, and career progression.

Qualifications

Required Skills and Qualifications

  • Kubernetes Certified Administrator (CKA) or equivalent experience in production environments. CKS advantage.
  • NVIDIA GPU-Accelerated server platform knowledge
  • Data Engineering, Data Transformation, Data Migration tools knowledge

 

  • Knowledge of Nvidia AI software stack related to GPU orchestration
  • GPU based Cloud platform software stack knowledge incl. its dependencies on below layers
  • Strong experience with CI/CD tools (Jenkins, GitLab) and GitOps practices.
  • Proficiency with Helm charts and Kubernetes resource management.
  • Scripting/programming in Python or Bash; Infrastructure-as-Code with Terraform and Ansible.
  • Experience with container images (Docker, Podman) and image scanning.
  • Familiarity with object storage systems and persistent volume management.
  • Knowledge of monitoring and observability tools (Prometheus, Grafana).
  • Understanding of running AI/HPC workloads at scale.
  • Strong troubleshooting and operational support skills in mission-critical environments.

Preferred Attributes

  • Ability to automate repetitive operational tasks and build self-service capabilities for developers.
  • Security-conscious attitude in day-to-day operations.
  • Excellent communication and cross-team coordination skills.

 

    Additional Information

    You will be working in the European Union to meet our customers' data security and privacy requirements.

    * Please be informed that our remote working possibility is only available within Hungary due to European taxation regulation.

    Apply directly at Deutsche Telekom AG
    Location
    Budapest, Budapest, Hungary
    Work mode
    Hybrid
    Track
    IC (Individual Contributor)
    Type
    FULL_TIME
    PostedJun 5
    Deutsche Telekom AG logo
    Deutsche Telekom AG
    telekom.com
    View all 95 Deutsche Telekom AG roles →

    Was this listing helpful?