Staff Software Engineer - Databases SRE
Grafana Labs- Base salary
- £104k–£125k Published base salary range
- Location
- Remote - UK, Sweden, Spain or Germany Remote eligibility
- Employment
- Full-time Staff / Principal
About the job
Grafana Labs, the company behind the open observability cloud, is a 100% remote company with 1,600+ team members across 40+ countries. We are looking for a Staff Software Engineer - SRE to help support our highest value Grafana Cloud customers by increasing the reliability of our Cloud databases that are based on Mimir, Loki, Tempo, and Pyroscope. These databases are provided as a SaaS product from AWS, GCP, and Azure across all regions. The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers.
In this role, you will
- Partner closely with product engineering squads (embedded model)
- Own production reliability for high-SLA and complex customer environments
- Design and implement automation to scale our reliability practices
- Ensure our customers meet our SLO targets
- Define and evolve per-tenant SLOs and reliability models
- Proactively reduce SLO burn to prevent repeat incidents
- Serve as a primary escalation point and on-call for relevant incidents
- Lead customer-impacting incident response and post-incident reviews
- Contribute to design docs and code reviews
- Influence feature design to ensure production scalability and operability
- Build automation to eliminate toil where needed
- Improve alert quality and reduce noisy escalations
What we seek
- 8+ years engineering experience, 4+ in SRE/CRE/production engineering
- Strong preference for formal customer reliability engineering experience
- Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
- Strong experience with technical leadership, leading a team through projects, mentoring other engineers
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Experience with one or more programming languages (e.g. Go, Python, Java, etc.)
- Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
- Excellent problem-solving and troubleshooting skills
- Experience with blame-free Incident Response, writing high quality PIRs
- Ability to reason about performance, scaling, and failure modes
- Comfortable working within an engineering team with autonomy and self-direction
- Ability to partner deeply with product engineering teams
Compensation: In the UK, the Base compensation range for this role is £103,958 - £124,750. Actual compensation may vary based on level, experience, and skillset as assessed in the interview process.
Benefits: Benefits include equity, bonus (if applicable) and other benefits listed here.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.