This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Production Engineer based in Australia.
This is a senior engineering opportunity focused on building and operating resilient technology platforms at significant internet scale.
You’ll lead engineers using Site Reliability Engineering principles to improve reliability, performance, security, and scalability.
The role spans hybrid cloud infrastructure, containerized workloads, serverless technologies, automation, and distributed systems.
You’ll play a key role in solving complex production challenges and turning operational insights into lasting platform improvements.
Working across engineering and product teams, you’ll connect technical execution with broader business and customer objectives.
You’ll also mentor engineers and help cultivate a culture that balances rapid delivery with exceptional operational standards.
The environment is collaborative, fast-paced, and well suited to an analytical engineer who enjoys ownership and complex problem-solving.
Accountabilities
- Lead and support an engineering culture focused on shipping effectively while maintaining high standards for system stability, performance, security, and scalability.
- Manage and improve production infrastructure across public cloud environments, including AWS and GCP, through monitoring, alerting, investigation, and resolution of operational issues.
- Identify opportunities to automate repetitive operational and development tasks, reducing toil and creating scalable, efficient solutions.
- Serve as a technical point of contact during critical incidents, coordinating communication across teams and driving effective incident resolution.
- Analyze alert and incident trends, conduct root cause analysis, and implement structural improvements that prevent recurring issues.
- Execute production changes, security requests, deployments, and patching activities with a strong focus on reliability and risk management.
- Design, build, and maintain complex distributed systems, microservices, and self-healing infrastructure capable of operating at large scale.
- Partner with Product Management and Engineering leadership to align technical solutions, priorities, and execution with business and product roadmaps.
- Coach, mentor, and develop engineers while contributing to a culture of continuous learning and technical excellence.
Requirements
- 5+ years of professional experience in software engineering, DevOps, Site Reliability Engineering, or a closely related discipline, with experience operating production systems at internet scale.
- Strong background in system architecture, distributed systems, microservices, and self-healing infrastructure, with an understanding of efficient resource and network utilization.
- Advanced experience with Kubernetes and Docker for deploying and operating containerized workloads in public cloud environments.
- Strong programming and automation skills, preferably with Python and Go, as well as experience with object-oriented programming and scripting.
- In-depth knowledge of modern web-serving infrastructure, including Linux, Nginx and/or Apache, MySQL, PHP, and infrastructure automation tools.
- Hands-on experience with Infrastructure as Code and automation technologies such as Ansible and Terraform.
- Practical experience implementing observability solutions, including metrics, monitoring, logging, and alerting platforms such as Prometheus or EFK stacks.
- Strong analytical, troubleshooting, incident management, and root cause analysis capabilities.
- Ability to work collaboratively across technical and business teams and communicate complex technical topics effectively.
- Forward-thinking, proactive mindset with a strong interest in continuous learning and improving systems and engineering practices.
Benefits
- Competitive compensation appropriate to the Australian market and level of experience.
- Company stock options, providing employees with an ownership stake.
- Superannuation program.
- Employee Assistance Program.
- Supplemental maternity and paternity pay.
- Generous vacation allowance.
- One-time A$745 home office stipend.
- Company wellness days.
- Flexible work-from-home arrangements.
- Inclusive and diverse workplace focused on professional growth, collaboration, and continuous development.
- Opportunities to work with large-scale cloud infrastructure and complex distributed systems.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1