Senior Site Reliability Engineer
Apply now!
Candidate data
Senior Site Reliability Engineer
About the Role
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve highly available, production-grade cloud and Kubernetes environments. In this role, you will champion SRE best practices, build secure and scalable CI/CD pipelines, and support our data and platform engineering ecosystem.
You will work closely with software, data, and platform engineering teams to improve reliability, automation, observability, security, and operational efficiency across our infrastructure.
Key Responsibilities
-
SRE Advocacy & Observability: Promote core SRE principles, including SLOs, monitoring, alerting, logging, tracing, toil reduction, incident response, and runbook creation. Design and maintain comprehensive observability solutions using tools such as New Relic, PagerDuty, and Prometheus/Grafana.
-
Cloud & Kubernetes Management: Administer and maintain highly available Kubernetes clusters, including GKE and on-premise environments such as Kubespray. Manage cloud resources across GCP and AWS, including services such as Cloud SQL, Pub/Sub, and AWS EMR.
-
Infrastructure as Code (IaC): Provision, manage, automate, and scale cloud and on-premise infrastructure using Terraform and Terragrunt. Establish infrastructure standards that promote consistency, reliability, and repeatability.
-
CI/CD & DevSecOps: Build, secure, maintain, and optimize automated CI/CD workflows using platforms such as GitHub Actions, Concourse, and Bitbucket Pipelines. Integrate security and quality controls throughout the software delivery lifecycle.
-
Data & Platform Engineering Support: Support data platform technologies including Kafka, Airflow, and Hadoop. Leverage modern Kubernetes operators and automation frameworks to facilitate cluster migrations, deployments, and platform operations.
-
Incident Management & Reliability: Participate in incident response, troubleshoot complex production issues, conduct root-cause analysis, and implement long-term improvements to prevent recurring incidents.
-
Automation & Continuous Improvement: Identify operational bottlenecks and opportunities for automation, reducing manual processes and engineering toil while improving platform scalability and reliability.
-
Engineering Collaboration: Partner closely with software, data, and platform engineers to establish reliable deployment practices, improve system performance, and promote operational excellence.
Requirements & Qualifications
-
Experience: Proven experience working in a Senior SRE, DevOps, Platform Engineering, or similar role supporting production environments.
-
Kubernetes & Cloud Native: Strong hands-on experience with Kubernetes, including GKE and/or on-premise deployments, as well as cloud-native technologies such as ArgoCD, Istio, and Helm.
-
Infrastructure as Code & Automation: Proficiency with Terraform and Terragrunt, along with configuration management and automation tools such as Ansible or Salt.
-
CI/CD Expertise: Demonstrated experience designing, implementing, and optimizing CI/CD pipelines using GitHub Actions, Bitbucket Pipelines, Concourse, or comparable technologies.
-
Observability & Incident Response: Practical experience with application performance monitoring, metrics, distributed tracing, logging, alerting, and incident-management integrations using tools such as New Relic, PagerDuty, and Prometheus/Grafana.
-
Cloud & Data Ecosystem: Hands-on experience with GCP and/or AWS, with familiarity supporting data infrastructure technologies such as Kafka and Airflow.
-
Linux & Systems Engineering: Strong understanding of Linux systems, networking, security, performance troubleshooting, and production infrastructure.
-
Communication & Collaboration: Excellent written and spoken English, with the ability to communicate complex technical concepts clearly and collaborate effectively with software and engineering teams.
-
Mentoring: Ability and willingness to mentor engineers, share best practices, and contribute to a strong culture of reliability and technical excellence.
What You’ll Do
You will play a key role in ensuring the reliability, scalability, security, and operational excellence of our production infrastructure. You’ll have the opportunity to work across cloud infrastructure, Kubernetes, CI/CD, observability, and data platforms while helping shape the engineering practices that support our growing technology ecosystem.
Over 60% of our candidates get invited to an interview with our Clients.
Apply with the form below and we will reach out to you in the next 24h