Skip to main content

Senior Site Reliability Engineer

Honeycomb
Remote - United KingdomUpdated 3d ago
Base salary
£128k–£150k
Published base salary range
Location
Remote - United Kingdom
Remote eligibility
Employment
Full-time
Senior
Role family
Infrastructure
Developer tools
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

Honeycomb is an observability platform for teams who build and manage software that matters. We are a fully distributed company, and we invest in our people and care about how you orient to our culture and processes.

Honeycomb's Site Reliability Engineering (SRE) team works at the intersection of infrastructure, developer experience, and organizational enablement. We lead technically complex, cross-team projects that improve reliability, scale systems, and make life easier for engineering teams. Our work spans AWS infrastructure, Kubernetes, Helm, Terraform, Kafka, and other tools.

What you'll do

  • Help Honeycomb scale our backend systems to support our highest-volume customers.
  • Build organizational trust through transparent communication, giving and receiving direct and kind feedback.
  • Work with other backend teams to dive deep into our stack to make sure we’re getting the most out of our infrastructure.
  • Be trained, become, and then train others as an Incident Commander.
  • Help SRE and Honeycomb develop a healthy cross-Atlantic engineering culture.
  • Participate in the team’s on-call rotation as the EU side of a new follow-the-sun rotation.
  • Help the organization navigate tradeoffs between reliability and its other goals and priorities.
  • Optional: act as an external ambassador through blog posts, conference talks, and presentations.

What you'll bring

  • Strong experience in AWS and Kubernetes.
  • Experience performing cost analysis and reduction.
  • Solid Helm, Terraform, and CI/CD experience.
  • Project management skills.
  • Software engineering experience (Golang is a plus, and so is performance engineering).
  • Experience with Kafka or another high-volume distributed system.
  • Excellent written and spoken communication skills.
  • A curiosity to learn how people and systems work.
  • Familiarity with observability concepts (SLOs, instrumentation) and data-driven decision making.
  • Comfort operating in ambiguity, with a bias for action and experimentation.
  • Interest in both the technical and human sides of reliability engineering.
  • Experience working in geographically distributed teams.

Compensation: Base Salary based on level of experience £127,670—£150,200 GBP.

Benefits: Generous equity, unlimited PTO, home office/co-working/internet stipend, full benefits coverage, up to 16 weeks of paid parental leave, annual development allowance.

Application: Please apply via the provided link. Note that we cannot currently sponsor or support visa transfers at this time.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.