Booking NL
Site Reliability Engineer (For independent contractors)
About the team
The BKS (Booking Kubernetes Service) team builds and operates the internal Kubernetes platform underpinning hundreds of engineering teams at Booking.com. We run a large fleet of EKS and on-premises clusters across multiple regions. We try to treat operations as a software problem, when something is painful, we want to automate it; when something breaks, we learn from it.
The role
This is a hands-on operational SRE role focused on keeping the BKS platform healthy, up-to-date, and well-supported. You will spend most of your time on cluster maintenance, component upgrades, and helping internal engineering teams run their workloads successfully on BKS. You may also act as a consultant to product teams on Kubernetes best practices and reliability topics.
Day-to-day responsibilities
Cluster operations & maintenance
Plan and execute Kubernetes version upgrades across EKS and on-premises clusters, coordinating with internal teams to minimise disruption
Perform routine maintenance: add-on upgrades, storage and networking configuration, component upgrades for monitoring, security and other tooling
Act on cluster health issues across the fleet and proactively work on degradation signals before they become incidents
Internal customer support
Be the first point of contact for engineering teams running workloads on BKS: triaging issues, diagnosing failures, and guiding teams to resolution
Help teams understand platform capabilities, quota management, and best practices for running reliable workloads. Evaluate quota requests and current usage requirements and weigh them against cluster capabilities.
Contribute to runbooks and FAQs so common questions are answered before they become support requests
Toil reduction & automation
Identify repetitive manual tasks and reduce them through scripting and automation
Flag and address technical debt that slows down operations or increases risk
Partner with the wider BKS team on tooling improvements that reduce the operational burden across the fleet
Required skills
Kubernetes operational experience: node management, upgrades, debugging workloads, cluster health
Experience with managed Kubernetes (EKS or equivalent) or on-premises cluster operations
Comfortable with observability tooling: Prometheus/VictoriaMetrics, Grafana, alerting pipelines
Strong written and verbal communication: clear in tickets, runbooks, and support threads
Able to work independently, prioritise effectively, and keep commitments
Nice to have
Terraform for infrastructure provisioning
Puppet or similar configuration management for on-premises environments
AWS experience
Experience supporting internal developer platforms or infrastructure teams
What we expect
Think Customer First: the engineering teams relying on BKS are your customers; their problems are your problems
Own It: take issues through to resolution, raise blockers early, and follow through without needing to be chased
Succeed Together: share knowledge, document what you learn, and make the team stronger for it
Learn Forever: whether it’s a new technology or a Booking-specific tool, knowing how to tackle something unknown by asking the right questions
Do The Right Thing: when something looks risky, an unsafe upgrade path, a coverage gap, a missed dependency, say so and help address it
What this role is not
A pure software development role, the focus is operational excellence and customer support, not feature building
A solo heroics role, we escalate, pair, and work in the open
A reactive-only role, proactive maintenance and toil reduction are just as important as incident response
Booking.com is an equal opportunity employer. We value diverse perspectives and are committed to building an inclusive environment where everyone can do their best work.