End Date
Sunday 19 July 2026We Support Flexible Working – Click here for more information on flexible working options
Flexible Working Options
Hybrid WorkingJob Description Summary
Our Site Reliability Engineering (SRE) team plays a key role in improving the reliability, resilience and operational excellence of the Analytics & AI Platform.Job Description
What you’ll do
Design, implement and continuously improve monitoring, alerting and observability solutions that enable reliable production services.
Define, measure and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Error Budgets to drive service reliability.
Work collaboratively with Production Support and engineering teams to investigate complex production incidents and support service restoration.
Support Post Incident Reviews (PIRs) and Problem Management activities by identifying reliability improvements and driving preventative engineering actions.
Reduce operational toil through automation, tooling and engineering improvements, enabling teams to focus on higher-value engineering work.
Work closely with Production Support teams to develop and continuously improve operational runbooks, support processes and service readiness.
Partner with engineering teams to embed reliability, resilience and operational best practices throughout the software development lifecycle.
Collaborate with Product and Engineering teams to improve service operability, ensuring applications are designed with appropriate monitoring, alerting, diagnostics and operational documentation before entering production support.
Identify reliability risks and proactively deliver engineering improvements that reduce incidents and improve service resilience.
Support the onboarding of new applications into the SRE operating model by ensuring agreed reliability, observability and operational standards are achieved.
Coach and mentor engineers, sharing knowledge and promoting engineering excellence across the platform.
Contribute to the evolution of SRE standards, tooling and engineering practices across the Analytics & AI Platform.
Where appropriate, provide line management, coaching and performance development for engineers within the team.
What you’ll need
Essential
Strong understanding of Site Reliability Engineering principles and practices.
Experience designing and implementing observability solutions, including monitoring, logging and alerting.
Experience defining and improving Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Error Budgets.
Strong troubleshooting and technical investigation skills across complex production environments.
Experience supporting incident management, Post Incident Reviews and Problem Management activities.
Experience identifying operational toil and delivering automation to improve reliability and efficiency.
Experience using Infrastructure as Code and CI/CD tooling.
Strong scripting or programming skills in one or more languages such as Python, Java, JavaScript, PowerShell or Bash.
Experience working with cloud technologies and modern application platforms.
Strong understanding of cloud security, networking and operational resilience.
Excellent stakeholder management, communication and collaboration skills.
Ability to work effectively across Product, Engineering and Production Support teams.
Desirable
Experience with Kubernetes and containerised platforms.
Experience with Azure, Google Cloud Platform or AWS.
Experience using observability platforms such as Dynatrace, Grafana, Prometheus, ELK or Splunk.
Experience working within Financial Services or another regulated industry.
Experience using Jira, Confluence and Agile delivery practices.
Experience mentoring or coaching engineers.
Previous line management experience or a desire to develop people leadership capability.
Relevant Cloud, DevOps or SRE certifications
Want to know if this job is worth applying to?
Hyderabad Knowledge Park Tower 2, India
Hybrid
