Want to know if this job is worth applying to?
North America Remote / San Francisco, CA
Hybrid
Location: North America Remote / San Francisco · Full-Time
Andromeda gives AI companies access to the kind of scaled compute once reserved for hyperscalers. Our platform connects 100+ AI customers to 50+ global providers, with billions of GPU-hours supported, and those numbers are all rapidly growing. We combine enterprise-grade reliability with the speed and economics of an open market, serving teams running everything from large-scale training to production inference.
Nat Friedman (former CEO of GitHub) and Daniel Gross (former head of AI at Apple, YC partner) started Andromeda in 2023 with a single GPU cluster. It filled almost immediately. Three years later, we're a $1.5B company, profitable since day one, with a Series A from Paradigm to scale the platform globally. The global flow of compute is already a multi-trillion dollar market, and our team is building the infrastructure that enables it to continue to scale.
The problem is deceptively hard. Not all compute is equal: interconnect, networking, OEM, firmware, and cluster age all vary across providers, and the differences matter at scale. Our platform benchmarks and validates capacity, takes positions, structures contracts, and operates clusters globally, delivering a consistent product regardless of where it runs. No one else has built this layer, and the AI industry can't scale without it.
As an Infrastructure Product Engineer, you will play a pivotal role in building the backbone of Andromeda’s platform. You'll transform complex, real-world infrastructure challenges into scalable product capabilities that benefit our customers.
Positioned at the intersection of infrastructure and product engineering, this role is deeply technical and systems-oriented, yet laser-focused on building solutions with broad leverage.
Design and develop core platform components, including infrastructure orchestration, provisioning, and lifecycle management solutions.
Build robust APIs, services, and control planes that abstract over diverse infrastructure types (VMs, Kubernetes, bare metal, schedulers).
Translate customer usage patterns into product requirements, delivering impactful features and improvements.
Create automation and internal tooling to eliminate manual or ad-hoc operational work.
Enhance reliability, performance, and observability at the platform level, emphasizing durable improvements over quick fixes.
Collaborate with peer teams to define clear ownership boundaries between platform capabilities and customer-specific solutions.
Write clean, maintainable, and well-documented code with a focus on long-term sustainability.
Participate in technical design discussions and contribute to the architectural evolution of our platform.
5+ years of experience in Infrastructure, Platform, or Backend Engineering roles.
Strong systems fundamentals: deep understanding of Linux, networking, storage, and distributed systems.
Proven expertise with Kubernetes, VMs, or bare-metal environments.
Advanced software engineering skills; capable of building production-grade APIs and services (Python, Go, or similar).
Extensive experience with infrastructure as code and automation tools (Terraform, Ansible, Helm, etc.).
Demonstrated ability to navigate ambiguity and distill complex problems into clear, maintainable abstractions.
Product-focused mindset: care about interfaces, defaults, reliability, and sustainable operations.
Excellent written and verbal communication skills; effective collaborator across engineering and product functions.
Nice to Have:
Hands-on experience with GPU or AI infrastructure.
Experience with control-plane or orchestration systems.
Background spanning both infrastructure and application/backend engineering.
Experience architecting multi-tenant systems.
Strong skills in technical writing and design documentation.
Early-stage startup experience.
This is a true builder’s opportunity: you’ll have ownership and autonomy to shape our systems, engage directly with customers and providers, and lay the foundations for scalable, reliable AI infrastructure. Join us at Andromeda and help power the future of AI.