This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in India.
This is a senior engineering role focused on building and operating reliable, scalable infrastructure and software for distributed compute environments. You will help solve complex reliability challenges through proactive troubleshooting, automation, observability, and systems engineering. The role combines infrastructure expertise with software development and close collaboration across operations, engineering, and support teams. You will design and maintain internal tooling and observability platforms that improve system visibility, performance, and service reliability. Working across a diverse technology landscape, you will contribute to modernizing existing systems while supporting the rollout of new applications and services. This opportunity is ideal for an experienced engineer who enjoys solving large-scale technical problems and improving the resilience of critical platforms.
Accountabilities:
- Solve complex infrastructure and reliability challenges through proactive troubleshooting, automation, systems programming, and structured root-cause analysis.
- Deploy, operate, and maintain observability platforms and internal engineering tools that improve system visibility, reliability, and operational performance.
- Partner with application development, operations, and support teams to ensure services are reliable, scalable, performant, and usable.
- Provide technical guidance to engineers and developers, helping teams establish confidence that their services meet reliability and performance expectations.
- Investigate and troubleshoot complex production issues across distributed systems, infrastructure, networking, and application environments.
- Contribute to the modernization of existing tooling and the development of new applications supporting compute and distributed infrastructure services.
- Improve operational processes through automation, infrastructure as code, monitoring, and repeatable engineering practices.
- Collaborate across teams to identify reliability risks, strengthen service-level performance, and continuously improve operational excellence.
Requirements:
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- 6+ years of experience in Site Reliability Engineering, infrastructure engineering, DevOps, systems engineering, or a closely related field.
- Strong Linux system administration expertise and practical experience troubleshooting production environments.
- Solid understanding of networking fundamentals, including TCP/IP, DNS, routing and switching, and storage concepts.
- Hands-on experience with containerized environments and Kubernetes, including operating, monitoring, and troubleshooting production systems.
- Strong knowledge of CI/CD and DevOps practices, with practical experience using tools such as Jenkins, Git, Prometheus, and Grafana.
- Experience implementing Infrastructure as Code and configuration management using Terraform, Ansible, SaltStack, or comparable technologies.
- Familiarity with cloud storage systems and distributed infrastructure environments.
- Strong automation and scripting capabilities using Python, Bash, Go, Rust, or similar programming languages.
- Strong analytical and problem-solving skills, with the ability to investigate complex technical issues and develop reliable long-term solutions.
- Excellent collaboration and communication skills, with the ability to work effectively across engineering, operations, and support teams.
- Comfortable working in a fast-paced environment where systems and technologies continuously evolve.
Benefits:
- Remote work opportunity in India, with flexibility to work from home, an office, or a combination depending on role requirements.
- Opportunity to work on large-scale distributed systems and complex Site Reliability Engineering challenges.
- Exposure to modern technologies across cloud, compute, Kubernetes, observability, automation, and infrastructure engineering.
- Collaboration with highly skilled engineering, operations, and application development teams.
- Opportunity to influence system reliability, scalability, monitoring, and operational excellence across critical services.
- Support for employee health, well-being, financial security, and life beyond work through comprehensive benefits.
- Professional growth through exposure to diverse technologies, modernization initiatives, and large-scale engineering environments.
- Flexible workplace approach designed to support productive and effective ways of working.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1