Lead Site Reliability Engineer

Lead • Full-time • Remote în Romania

HDAVL001
1
Remote

Descriere

What you'll do

You will lead how reliability is engineered across Partenerul nostru's global SaaS platform as it scales and moves toward an AI-first operating model. This role focuses on building a modern, automation-first reliability ecosystem that improves system stability, reduces operational risk and enables faster, safer product delivery.

You will work across multi-cloud environments to design self-healing systems, advance observability and modernise deployment practices. As a senior individual contributor, you will also raise the technical bar by shaping standards, mentoring engineers and driving measurable improvements in reliability and performance.

Responsibilities

  • Own and evolve the reliability strategy for distributed SaaS systems across multi-cloud platforms.
  • Design and implement AI-driven operations, including predictive monitoring, anomaly detection and automated root cause analysis.
  • Build and scale observability solutions using tools such as Prometheus, Grafana and OpenTelemetry.
  • Create self-healing systems and automation frameworks that reduce manual operational work.
  • Improve deployment practices using feature flags, progressive delivery and safe rollout strategies.
  • Ensure reliability and performance of CI/CD pipelines and infrastructure as code environments.
  • Strengthen system availability, scalability and fault tolerance across Kubernetes-based platforms.
  • Lead incident response, improve recovery times and implement lasting fixes through post-incident reviews.
  • Integrate AI-driven workflows into incident detection, triage and resolution to improve operational efficiency.
  • Mentor engineers and drive adoption of automation-first and AI-first reliability practices.

What you'll need to be successful

  • 10+ years of experience in SaaS, distributed systems or site reliability engineering.
  • Strong programming skills in Go, Java or Python.
  • Deep experience with observability tools such as Prometheus, Grafana and OpenTelemetry.
  • Hands-on experience with Kubernetes, containerisation and multi-cloud platforms (AWS, GCP, Azure or OCI).
  • Strong understanding of Linux systems, networking and cloud-native architectures.
  • Proven ability to design automation, improve system reliability and apply AI or machine learning to operational workflows.

Partenerul nostru is an AI-first company

AI is embedded in Partenerul nostru's workflows, decision-making and products, and success in this role requires embracing AI as an essential capability. You will bring experience using AI and AI-related technologies, apply AI every day to business challenges to improve efficiency and drive results, and keep growing by staying curious about new trends and best practices and sharing what you learn so others can benefit too.

Ești interesat(ă)? Aplică acum!

Promitem să păstrăm datele tale personale în siguranță și vor fi folosite doar pentru a aplica la această poziție.

Câmpurile marcate cu * sunt obligatorii. Aplicând, ești de acord cu Politica De Confidențialitate și Termenii de Utilizare a site-ului.

Așteaptă, avem mai multe...

Trebuie să fie un job perfect pentru tine, iată câteva joburi similare.

Senior Golang Engineer
Senior level • Full-time • HDFAC007
Hibrid Cluj-Napoca
Senior Java Software Engineer
Senior level • Full-time • HDFAC006
Hibrid Cluj-Napoca
Java Technical Lead
Lead • Full-time • HDFAC005
Hibrid Cluj-Napoca

Îți prezentăm consola
developerului.

Înscrie-te la newsletter-ul nostru și vei primi actualizări periodice cu postări noi, concursuri, evenimente și oportunități de locuri de muncă.

$