Skip to main content

Staff Software Engineer - Databases SRE

Grafana Labs
Remote - UK, Sweden, Spain or GermanyUpdated 5d ago
Base salary
£104k–£125k
Published base salary range
Location
Remote - UK, Sweden, Spain or Germany
Remote eligibility
Employment
Full-time
Staff / Principal
Role family
Engineering
Developer tools
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

Grafana Labs, the company behind the open observability cloud, is a 100% remote company with 1,600+ team members across 40+ countries. We are looking for a Staff Software Engineer - SRE to help support our highest value Grafana Cloud customers by increasing the reliability of our Cloud databases that are based on Mimir, Loki, Tempo, and Pyroscope. These databases are provided as a SaaS product from AWS, GCP, and Azure across all regions. The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers.

In this role, you will

  • Partner closely with product engineering squads (embedded model)
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale our reliability practices
  • Ensure our customers meet our SLO targets
  • Define and evolve per-tenant SLOs and reliability models
  • Proactively reduce SLO burn to prevent repeat incidents
  • Serve as a primary escalation point and on-call for relevant incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design docs and code reviews
  • Influence feature design to ensure production scalability and operability
  • Build automation to eliminate toil where needed
  • Improve alert quality and reduce noisy escalations

What we seek

  • 8+ years engineering experience, 4+ in SRE/CRE/production engineering
  • Strong preference for formal customer reliability engineering experience
  • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
  • Strong experience with technical leadership, leading a team through projects, mentoring other engineers
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Experience with one or more programming languages (e.g. Go, Python, Java, etc.)
  • Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting skills
  • Experience with blame-free Incident Response, writing high quality PIRs
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working within an engineering team with autonomy and self-direction
  • Ability to partner deeply with product engineering teams

Compensation: In the UK, the Base compensation range for this role is £103,958 - £124,750. Actual compensation may vary based on level, experience, and skillset as assessed in the interview process.

Benefits: Benefits include equity, bonus (if applicable) and other benefits listed here.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.