This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead - GPU Infrastructure based in Turkey.
This is a hands-on technical leadership role responsible for architecting and delivering a large-scale GPU infrastructure platform.
You will lead the evolution from managed Kubernetes workloads toward bare-metal GPU infrastructure, including Slurm-based research computing and Kubernetes-powered inference.
The role combines deep systems expertise with engineering leadership, team management, and direct ownership of architecture and delivery.
You will oversee a distributed team spanning backend, frontend, DevOps, QA, and documentation while maintaining high technical standards.
Your work will support research, model training, and managed inference workloads requiring reliable, scalable, and observable GPU compute.
You will also serve as the primary technical interface with infrastructure partners, hardware providers, and internal platform consumers.
This is an opportunity to shape the architecture and operational foundations of a sophisticated GPU platform in a fully remote environment.
Accountabilities:
The Technical Lead will own the platform architecture and engineering delivery while remaining deeply involved in technical decisions, infrastructure operations, team leadership, and partner relationships.
- Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and ongoing architecture documentation.
- Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation.
- Establish engineering standards, oversee code and design reviews, manage release gates, conduct one-to-ones, and provide growth and performance feedback.
- Design, build, and operate a managed Slurm service supporting research and model-training workloads.
- Own Slurm controllers, accounting, partitions, login nodes, node onboarding, acceptance testing, driver and CUDA baselines, upgrades, stalled-job detection, node health, draining, autohealing, storage visibility, identity, and workload isolation.
- Lead GPU infrastructure operations on bare-metal environments, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
- Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator.
- Oversee GPU isolation using technologies such as KubeVirt and VFIO and manage day-two infrastructure operations, upgrades, backup, recovery, and node replacement.
- Define managed inference architecture covering serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute capabilities.
- Establish observability across the control plane, GPU fleet, and application layers through metrics, logging, alerting, and SLOs.
- Lead incident response, post-incident reviews, and the development of an on-call model that is sustainable for a lean engineering organization.
- Act as the primary technical interface with infrastructure partners and vendors, translating requirements into written specifications and acceptance tests.
- Manage technical escalations with partners through resolution and contribute to capacity planning and hardware sourcing decisions.
- Work directly with research, model-training, and product teams to translate workloads into platform requirements and manage capacity constraints.
- Hire and develop members of the platform team while maintaining a high technical bar.
- Contribute to architecture decisions involving distributed systems, high-performance computing, networking, storage, virtualization, and GPU workloads.
Requirements:
The ideal candidate combines deep hands-on GPU and infrastructure expertise with proven technical leadership experience. They should be comfortable operating complex production systems, making architecture decisions, leading distributed teams, and remaining close to the code and infrastructure.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1