This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE and DevOps Team Manager based in India.
This is a hands-on leadership role responsible for the reliability, scalability, automation, and operational maturity of a growing SaaS platform. You will lead and develop an SRE/DevOps team while staying technically engaged with the systems and challenges they support. The role combines people leadership with ownership of incident management, cloud infrastructure, observability, and reliability practices. You’ll work closely with engineering leadership to align infrastructure priorities with business, security, compliance, and cost objectives. As part of a globally distributed environment, you’ll coordinate operations across Indian Standard Time and US time zones. This is an opportunity to establish strong SRE practices and shape a culture built around ownership, resilience, and continuous improvement.
Accountabilities:
- Lead, mentor, and develop a team of SRE and DevOps engineers, working with engineering leadership to assess capabilities, define team needs, and support individual growth.
- Own the end-to-end incident management process, including on-call rotations, escalation procedures, incident command, severity frameworks, blameless postmortems, root cause analysis, and preventative actions.
- Establish and promote SRE principles across services, including SLIs, SLOs, error budgets, capacity planning, reliability standards, and observability practices.
- Drive operational excellence by improving processes, automation, reliability, security, and cost efficiency while balancing operational priorities with feature delivery.
- Coordinate a distributed support and engineering model across IST and US time zones, ensuring effective handoffs, coverage, communication, and escalation.
- Partner with engineering leadership to prioritize infrastructure investments and align platform initiatives with business, security, compliance, and operational objectives.
- Monitor and report on reliability metrics, team health, operational performance, incident trends, and progress against reliability goals.
- Guide day-to-day technical priorities and remain hands-on enough to support the team through complex infrastructure and operational challenges.
- Contribute to roadmap planning, department-wide initiatives, and the continuous improvement of the SRE/DevOps function.
- Promote a culture of accountability, ownership, continuous learning, and proactive reliability engineering across the team.
Requirements:
- 2–3 years of experience as a Team Lead or Manager within an SRE, DevOps, Platform Engineering, or similar infrastructure environment.
- Proven experience managing and coordinating distributed teams across IST and US time zones, including on-call coverage, operational handoffs, and cross-regional collaboration.
- Strong experience managing incident response processes, including on-call operations, escalation paths, incident command, severity frameworks, blameless postmortems, and RCA follow-through.
- Strong hands-on expertise with cloud infrastructure and multi-cloud environments, including AWS, Kubernetes, GKE, EKS, managed databases, object storage, networking, and IAM.
- Solid experience with Infrastructure-as-Code technologies such as Terraform, Pulumi, or equivalent tools.
- Experience with modern CI/CD platforms and practices, including GitHub Actions, Jenkins, or similar technologies.
- Strong understanding of observability across metrics, logging, tracing, monitoring, and alerting.
- Experience establishing or scaling SRE/DevOps functions, on-call programs, reliability practices, or operational processes is highly desirable.
- Experience in FinTech or another regulated environment, particularly where security, risk, and compliance requirements are important, is a plus.
- Experience improving database and data infrastructure operations, as well as observability across multiple cloud technologies, is advantageous.
- Strong understanding of designing cloud environments for cost efficiency, resilience, scalability, security, and high availability.
- Excellent leadership, communication, prioritization, and problem-solving skills, with the ability to operate effectively in a distributed and fast-moving environment.
- A hands-on mindset with the ability to balance technical depth, people leadership, strategic planning, and day-to-day operational demands.
Benefits:
- 100% remote position based in India.
- Opportunity to lead and develop an SRE/DevOps team within a high-growth SaaS environment.
- Hands-on exposure to multi-cloud infrastructure, Kubernetes, automation, observability, and modern reliability engineering practices.
- Opportunity to shape SRE processes, incident management, on-call practices, and operational standards.
- Collaboration with engineering leadership and globally distributed technical teams.
- Exposure to security, compliance, resilience, and cost optimization within a technology-driven financial services environment.
- Career development opportunities through technical leadership, team management, and cross-functional initiatives.
- Flexible collaboration across Indian and US time zones within a distributed working environment.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1