Job Summary
What you'll do
You will lead how reliability is engineered across Our partner's global SaaS platform as it scales and moves toward an AI-first operating model. This role focuses on building a modern, automation-first reliability ecosystem that improves system stability, reduces operational risk and enables faster, safer product delivery.
You will work across multi-cloud environments to design self-healing systems, advance observability and modernise deployment practices. As a senior individual contributor, you will also raise the technical bar by shaping standards, mentoring engineers and driving measurable improvements in reliability and performance.
Responsibilities
- Own and evolve the reliability strategy for distributed SaaS systems across multi-cloud platforms.
- Design and implement AI-driven operations, including predictive monitoring, anomaly detection and automated root cause analysis.
- Build and scale observability solutions using tools such as Prometheus, Grafana and OpenTelemetry.
- Create self-healing systems and automation frameworks that reduce manual operational work.
- Improve deployment practices using feature flags, progressive delivery and safe rollout strategies.
- Ensure reliability and performance of CI/CD pipelines and infrastructure as code environments.
- Strengthen system availability, scalability and fault tolerance across Kubernetes-based platforms.
- Lead incident response, improve recovery times and implement lasting fixes through post-incident reviews.
- Integrate AI-driven workflows into incident detection, triage and resolution to improve operational efficiency.
- Mentor engineers and drive adoption of automation-first and AI-first reliability practices.
What you'll need to be successful
- 10+ years of experience in SaaS, distributed systems or site reliability engineering.
- Strong programming skills in Go, Java or Python.
- Deep experience with observability tools such as Prometheus, Grafana and OpenTelemetry.
- Hands-on experience with Kubernetes, containerisation and multi-cloud platforms (AWS, GCP, Azure or OCI).
- Strong understanding of Linux systems, networking and cloud-native architectures.
- Proven ability to design automation, improve system reliability and apply AI or machine learning to operational workflows.
Our partner is an AI-first company
AI is embedded in Our partner's workflows, decision-making and products, and success in this role requires embracing AI as an essential capability. You will bring experience using AI and AI-related technologies, apply AI every day to business challenges to improve efficiency and drive results, and keep growing by staying curious about new trends and best practices and sharing what you learn so others can benefit too.