All jobs

grafanalabs

Staff Software Engineer - Databases SRE | UK | Remote

Location flexible

About this role

Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings AI to observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. We are a 100% remote company with team members across 40+ countries, backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital. Learn more at grafana.com and follow us on LinkedIn and X . We’re scaling fast and staying true to what makes us different: an open-source legacy, a global collaborative culture, and a passion for meaningful work. Our team thrives in an innovation-driven environment where transparency, autonomy, and trust fuel everything we do. You may not meet every requirement, and that’s okay. If this role excites you, we’d love you to raise your hand for what could be a truly career-defining opportunity. This is a remote opportunity and we are looking for candidates from the UK, Sweden, Spain or Germany. About the role: We are looking for a Staff Software Engineer - SRE to help us support our highest value Grafana Cloud customers by increasing the reliability of our Cloud databases that are based on Mimir, Loki, Tempo, and Pyroscope. We provide these databases as a SaaS product from AWS, GCP, and Azure across all regions. The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will: Partner closely with product engineering squads (embedded model) Own production reliability for high-SLA and complex customer environments Design and implement automation to scale our reliability practices Ensuring our customers meet our SLO targets Define and evolve per-tenant SLOs and reliability models Proactively reduce SLO burn to prevent repeat incidents Serving as a primary escalation point and on-call for relevant incidents Lead customer-impacting incident response and post-incident reviews Contribute to design docs and code reviews Influence feature design to ensure production scalability and operability Build automation to eliminate toil where needed Improve alert quality and reduce noisy escalations We seek a staff software engineer operating at the intersection of customer needs, production systems, and product engineering. We invest heavily in developer productivity. You can use modern AI coding assistants as part of your daily workflow (your choice of tools, within security guidelines), backed by a company-funded usage budget so you can iterate quickly without unnecessary friction. We encourage pragmatic AI-assisted development: faster prototyping, test generation, refactors, documentation, and incident follow-ups—always paired with strong code review and quality standards. You’ll also have access to frontier models (e.g., GPT-Codex 5/3, Claude Opus 4.6, Gemini 3 Pro). What we seek: 8+ years engineering experience, 4+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience. Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.). Strong experience with technical leadership, leading a team through projects, mentoring other engineers on the team and serving as a force-multiplier Experience operating multi-tenant systems in production Strong experience designing and implementing SLOs Experience with one or more

Skills and categories

R&D : Databases

Listing provided by Arbeitnow. Verify availability on the source board before applying — Kartavyam aggregates public listings as they appear and does not manage this application.

Similar roles

Browse all jobs