
Want to know if this job is worth applying to?
Greater London, England, United Kingdom
Remote only
WHO WE ARE
Enterprise work still moves by hand: copy-pasting between spreadsheets, endless email threads, and clunky legacy UIs. We started Duvo to end that for good.
We've already earned the trust of a range of customers, and our agents are helping them automate business-critical processes. We are growing fast, but to win from here we need exceptional people; that's where you come in.
WHAT WE ARE BUILDING
We're building the AI operations platform for large enterprises, currently focused on retail and consumer packaged goods customers. In Duvo, customers build AI agents that execute work wherever it needs to get done—SAP, spreadsheets, supplier portals, email, APIs, you name it. Duvo is heavy on browser and computer use.
In Duvo, business users specify the outcome; agents plan, act, request approvals on exceptions, and learn with every run. To help customers understand what to automate, we also help them map their processes by digesting interviews and internal documentation, which gets our foot in the door. We start with automating the parts of companies that we know best (category management, supply chain, finance ops) where we can show value fast, then expand to adjacent functions and sectors.
Velocity is our moat: ship fast, iterate faster, compound learning.
THE ROLE
You will lead the efforts to level up reliability across different facets of Duvo within a team whose mission is to earn and keep the trust of customers running their critical business workflows on Duvo.
First, as LLMs enable both developers and non-developers to ship code much more quickly, we need seasoned SREs who roll up their sleeves and evolve Duvo's observability, infrastructure topology, and release process to meet the challenges that high-velocity coding brings.
Second, at the core of Duvo, we have agents executing complex customer workflows through a combination of reasoning and heavy tool use. You'll ensure we spot misbehaving agents early, have a good understanding of their failure modes, and identify the causes of high latency.
When bugs creep in, the tooling you've built sounds the alarm before customers notice and gives the on-call team the tools to assess impact and mitigate issues in minutes, not hours. You're not afraid to introduce friction to ensure that post-mortem action items are executed across the engineering org in a timely fashion.
WHAT WE'RE LOOKING FOR
These are non-negotiables—the things we'll specifically evaluate you on:
- Distributed systems experience. You've designed and operated systems that scale. You understand failure modes, capacity planning, and the trade-offs between consistency, availability, and latency in real production environments.
- Security mindset. You'll handle enterprise data flowing through sandboxed environments, manage KMS encryption, configure Cloud Armor WAF rules, and ensure network isolation between tenant workloads. Security is a default consideration, not an afterthought.
- Observability and incident response. You build monitoring and alerting that catches problems before customers do. When incidents happen, you lead structured responses, find root causes, and drive lasting fixes — not just restarts.
- Infrastructure as code and automation. You automate everything you can. You've worked with IaC tools, CI/CD pipelines, and container orchestration in production. Manual runbooks make you uncomfortable.
- Shipping and ownership. You don't just maintain systems — you improve them. You take ownership of reliability projects from proposal to production, and you measure the results.
- Judgment on where to invest. You'll decide what to automate first, where to invest in reliability vs. ship speed, and make incident calls with incomplete information.
YOU MIGHT ALSO
- Have experience with GCP, Kubernetes, or similar cloud-native infrastructure.
- Have worked with sandboxed execution environments or multi-tenant isolation.
- Be comfortable with AI/ML production systems, understanding the unique reliability challenges of LLM-based applications.
- Have a product engineering background — you've built features and understand the developer experience you're supporting.
THIS IS NOT FOR YOU IF
- You want a traditional ops role, where you follow runbooks — we're building the reliability practice, not maintaining one.
- You want to focus on building AI features — see AI Platform Engineer if that's more the area of interest for you.
LOCATION
We're looking for people within two hours of Prague (CET), so the team overlaps for most of the working day.
OUR TECH STACK
- GCP (Cloud Run, GKE, GCS)
- Terraform, Docker
- Datadog
- TypeScript and Python services (you'll read and occasionally modify application code, but deep language expertise isn't required)
- Postgres, Redis
HOW WE WORK
These are real tradeoffs we've made, not aspirations:
- Initiative-driven. We organize around customer problems, not org charts. Problems surface through product feedback, competitive analysis, and direct customer conversations — then we prioritize, build, and ship weekly.
- Customer-obsessed. We solve real problems, not hypothetical ones. Features that don't move customer metrics get cut.
- Iterative by default. We ship small, learn fast, and never get attached to yesterday's code. This means things break sometimes — we fix forward.
- AI-first leverage. We use AI to move faster and focus human time where it matters most. If a tool can do it, a person shouldn't.
- Direct feedback. We give each other actionable feedback immediately. This can feel uncomfortable — we think that's worth it.
- Autonomy with accountability. We trust people to make decisions and hold them to outcomes, not process.
WHAT WE OFFER
- Unlimited AI budget. We don't just allow AI tools — we strongly encourage them. Want to try a new tool? Buy it. Want to automate part of your workflow? Do it.
- Autonomy to do your best work. Want to meet someone to learn from? Set it up. Want a mentor? Go get one. Want to fly out to talk to an important customer? Just ask.
- A real AI product with real customers. You're not building demos or internal tools. Enterprise customers use what you ship, and their feedback drives what you build next.
- A sharp, motivated team that values ownership and candor.
HOW WE HIRE
We respect your time and aim to move fast:
1. Discovery call (online, 30 min). We'll talk about you, how you think and whether there's mutual fit.
2. Technical interview (online, ~1 hour). Meet the team. We'll go deeper on your experience, system design, product thinking, and collaboration. No trick questions — we want to see how you think and build.
3. On-site trial day (1 day). Ship something small to production with us and see how we work together. Compensated.



