This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Expert Observability Engineer based in Brazil.
This role provides senior technical leadership across enterprise observability, reliability engineering, and cloud-native monitoring environments.
You will design and govern unified observability strategies spanning metrics, logs, traces, and events across complex technology landscapes.
The position combines architecture, automation, SRE practices, incident management, and platform engineering to improve service reliability.
You will lead innovative implementations, modernize legacy monitoring capabilities, and establish secure, scalable, production-ready observability patterns.
The role also involves close collaboration with global engineering and operations teams during major incidents, transitions, and knowledge-transfer programs.
Automation and infrastructure-as-code will be central to reducing operational effort, alert noise, and mean time to detect and resolve issues.
This is an opportunity to influence enterprise technology roadmaps while mentoring technical teams in a remote-first environment.
Accountabilities:
- Architect and govern a unified enterprise observability framework covering metrics, logs, traces, and events.
- Design observability solutions using platforms and technologies such as IBM Instana, Grafana, OpenTelemetry, Telegraf, InfluxDB, and Prometheus.
- Lead First-of-a-Kind (FOAK) implementations, evaluating emerging observability technologies and converting them into secure, repeatable production patterns.
- Define enterprise standards for telemetry pipelines, data retention, high-cardinality management, and observability cost optimization.
- Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Act as a senior technical escalation point for critical operational issues and major incidents.
- Lead P1/P2 incident war rooms and drive evidence-based Root Cause Analysis (RCA).
- Reduce Mean Time to Detect (MTTD), Mean Time to Resolve (MTTR), and alert noise through event correlation, dynamic thresholds, and dependency mapping.
- Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps practices.
- Automate deployment, configuration, monitoring, and self-healing workflows for agents and telemetry collectors.
- Integrate observability platforms with ITSM tools, ServiceNow, Netcool, and CI/CD pipelines.
- Design deep observability capabilities across Docker, Kubernetes, OpenShift, microservices, and multi-cloud environments.
- Correlate application performance telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies.
- Establish secure-by-design telemetry pipelines incorporating RBAC, TLS, secrets management, and image scanning.
- Lead complex knowledge-transfer programs, vendor transitions, and operational-readiness handovers for global 24x7 teams.
- Mentor cross-functional engineering and operations teams on observability, SRE, automation, and reliability practices.
- Contribute to enterprise technology roadmaps and influence architectural decisions across observability and platform engineering.
- Modernize legacy monitoring environments by introducing proactive, automated, and scalable observability capabilities.
Requirements:
- 12+ years of total IT experience, with at least 5–7 years operating as a Lead Architect, SRE, Principal Observability Engineer, or equivalent senior technical role in a large enterprise environment.
- Strong hands-on experience with observability and APM technologies, including IBM Instana, Grafana Enterprise and Alloy, Prometheus, OpenTelemetry, Telegraf, and InfluxDB.
- Experience with traditional monitoring platforms such as SolarWinds, Netcool, Elastic, and Splunk.
- Proven expertise with Kubernetes, Docker, OpenShift, and cloud environments across AWS, Azure, and/or GCP.
- Strong infrastructure knowledge across Linux/RHEL, Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.
- Hands-on experience with Ansible, Terraform, Python, Bash, webhooks, and CI/CD platforms such as GitHub Actions, GitLab, or Jenkins.
- Experience integrating monitoring and observability platforms with ServiceNow and other ITSM and operational tooling.
- Strong understanding of ITIL 4 principles and advanced Major Incident Management.
- Proven track record of migrating organizations from legacy monitoring solutions to proactive and automated observability platforms.
- Experience designing scalable and secure telemetry pipelines and time-series database environments.
- Strong understanding of SRE concepts, including SLIs, SLOs, error budgets, reliability engineering, and incident response.
- Experience implementing observability across distributed systems, microservices, Kubernetes, and multi-cloud architectures.
- Demonstrated experience leading FOAK technology implementations and complex vendor or operational transition programs.
- Strong automation mindset with the ability to turn operational processes into repeatable infrastructure and software workflows.
- Excellent troubleshooting, analytical, and root-cause investigation capabilities.
- Strong communication and stakeholder-management skills, with the ability to influence technical decisions across teams.
- Experience mentoring engineers and leading knowledge-transfer initiatives in global or distributed environments.
- CKA, cloud architecture certifications such as AWS or Azure, or relevant APM/observability certifications are preferred.
Benefits:
- CLT employment with a 40-hour weekly workload.
- Remote work in Brazil, with the flexibility to work from home when client-site presence is not required.
- SulAmérica Prestige health insurance, including coverage for eligible legal dependents.
- SulAmérica dental insurance, including coverage for eligible legal dependents.
- Prudential life insurance equivalent to 24x salary.
- MetLife private pension plan with company matching of up to 6%.
- Flash meal voucher and internet allowance totaling R$1,000 per month.
- Employee Assistance Program.
- Wellness program.
- Base salary plus additional compensation programs, subject to eligibility.
- Individual performance-based compensation opportunities, where applicable.
- Equity grant opportunities through the applicable Associate Equity Appreciation Program.
- Inclusive and diverse working environment with opportunities for professional development and collaboration.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1