This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer, AI Product based in India.
This role offers the opportunity to shape the engineering foundations behind AI-powered products operating at enterprise scale. You will design and evolve distributed systems capable of supporting high-volume AI agents while balancing reliability, performance, and cost. The position combines hands-on software engineering with technical leadership, mentorship, and architectural ownership. You will establish engineering practices around testing, observability, CI/CD, and production reliability that enable teams to move faster with confidence. Working closely with ML engineers and researchers, you will help turn advanced AI capabilities and agentic systems into dependable production experiences. With significant autonomy and broad technical influence, you will help define the platform architecture and engineering standards that support future AI product development.
Accountabilities:
- Architect and implement distributed systems capable of supporting AI agents and high-volume workloads while maintaining reliability, performance, scalability, and cost efficiency.
- Drive engineering excellence by establishing comprehensive testing strategies, observability practices, CI/CD pipelines, and operational standards that identify issues before they affect customers.
- Lead through technical expertise by mentoring engineers, sharing production engineering knowledge, and helping teams develop stronger approaches to designing and operating reliable systems.
- Own and evolve the technical roadmap for AI platform capabilities, making architectural decisions that support long-term scalability and maintainability.
- Collaborate closely with ML engineers and researchers to productionize advanced AI capabilities while maintaining system stability, operational reliability, and high engineering standards.
- Improve service reliability and operational maturity by identifying recurring failure patterns, reducing incidents, and establishing practices that support high availability.
- Investigate complex production issues, identify root causes, and implement durable solutions supported by appropriate monitoring, testing, and regression prevention.
- Make pragmatic architectural and engineering trade-offs, balancing technical quality, operational requirements, development speed, and the need to deliver production-ready solutions.
- Contribute hands-on code across the technology stack while providing technical direction for complex engineering initiatives.
- Help establish engineering standards and reusable platform capabilities that enable other teams to develop, test, deploy, and operate AI-powered products more efficiently and safely.
Requirements
- 8+ years of professional experience building and operating distributed systems in production environments.
- Demonstrated experience improving system reliability and taking production services toward high-availability targets, including 99.9%+ uptime.
- Deep knowledge of modern software engineering practices, including microservices, containerization, Kubernetes, and infrastructure as code.
- Strong programming and software design skills, with the ability to contribute hands-on across multiple layers of a modern technology stack.
- Proven experience leading significant technical initiatives and mentoring engineers or engineering teams.
- Strong production debugging and troubleshooting capabilities, including the ability to diagnose complex distributed-system failures and implement effective long-term solutions.
- Experience designing and implementing comprehensive observability strategies covering system health, application behavior, performance, and operational issues.
- Strong understanding of CI/CD, automated testing, deployment practices, and engineering processes that support reliable production delivery.
- Ability to make practical trade-offs between architectural perfection and delivering maintainable, production-ready solutions.
- Excellent technical communication and collaboration skills, with the ability to influence architecture and engineering practices across teams.
- Experience with LLM applications, AI/ML infrastructure, agent frameworks, prompt engineering, RAG architectures, or vector databases is advantageous.
- Familiarity with multi-model AI environments and platforms involving technologies such as OpenAI, Anthropic, Google, or comparable providers is a plus.
- Experience working in high-growth organizations or modern engineering environments with strong ownership and continuous improvement cultures is beneficial.
Benefits
- Opportunity to shape the engineering architecture behind sophisticated AI-powered products used by enterprise customers.
- Hands-on exposure to distributed systems, AI agents, LLM applications, modern infrastructure, and emerging AI engineering practices.
- Significant autonomy and technical influence over platform architecture, engineering standards, and long-term technology direction.
- Opportunity to work closely with ML engineers, researchers, and experienced software engineers on production AI systems.
- Leadership and mentorship opportunities with the ability to elevate engineering practices and help other developers build more reliable systems.
- Exposure to modern technologies and practices including Kubernetes, microservices, infrastructure as code, observability, automated testing, and CI/CD.
- Opportunity to solve complex scalability, reliability, and production engineering challenges at significant enterprise scale.
- An environment focused on engineering excellence, continuous improvement, collaboration, and pragmatic technical decision-making.
- Professional growth through ownership of high-impact technical initiatives and exposure to advanced AI product development.
- Inclusive workplace culture committed to equal opportunity, respect, collaboration, and supporting diverse perspectives.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1