
Monterail
DevOps / Senior Cloud Infrastructure Engineer - Freelancer
RemoteSeniorRemote eligibility not specifiedAI/MLInfrastructure
Our Take
Build and own the observability and incident management backbone for autonomous kitchen operations using AWS, Python, and Terraform.
What you’ll do
- Standardize Tech Ops Incident Management Platform using Jira Service Management
- Automate incident resolution workflows with AI tooling and runbooks
- Build fleet management and task automation for global operations
- Own and customize the data-driven observability layer via Grafana
What they’re looking for
- Senior experience as AWS Cloud or Site Reliability Engineer
- Fluent hands-on Python for tooling and automation
- Hands-on Terraform experience for Infrastructure as Code
- Proven track record owning production systems end-to-end
- Experience with Grafana dashboarding pipelines at scale
Skills & Focus Areas
- AWS
- Python
- Terraform
- Jira Service Management
- Grafana
- Observability
- Incident Management
As posted by Monterail
Job description
We are looking for a DevOps / Senior Cloud Infrastructure Engineer to join us on a freelance basis to build the observability and incident management backbone for autonomous kitchen operations.
4-5 month project | full time | remote
What we’re looking for
- Senior-level experience as an AWS Cloud Engineer or Site Reliability Engineer.
- Strong, demonstrable focus on metric-driven observability, monitoring, and alerting at scale.
- Fluent, hands-on experience in Python for tooling and automation.
- Hands-on experience with Terraform for Infrastructure as Code.
- A proven track record of designing, architecting, and owning production systems end-to-end.
- Ability to work completely independently without a detailed spec sheet or heavy direction.
- Experience with Jira Service Management or similar ITSM/incident platforms is a plus.
- Experience with Grafana dashboarding pipelines at scale is a plus.
- Exposure to AI-assisted ops tooling (AI-Ops, runbook automation) is a plus.
- Prior experience in a fast-moving hardware, robotics, or IoT fleet environment is a plus.
What you'll do
- Standardize and own the Tech Ops Incident Management Platform using Jira Service Management across our production fleets.
- Automate incident resolution workflows with AI tooling, including runbook generation and assignment.
- Design and implement proper documentation for all Tech Ops incident processes to ensure a clean handover.
- Build out fleet management and task automation to support global 24/7 remote operations.
- Own and customize the data-driven observability layer via Grafana across all internal tech teams.
- Work closely with key leadership stakeholders (Head of Infrastructure & Cloud, VP Engineering) to independently drive architecture decisions.
Requirements
What do we mean by freelance?
Read more at Monterail Tech Network