The Opportunity
Grafana Cloud is our composable observability platform that integrates metrics, logs, and traces with Grafana. It allows our customers to leverage the best open source observability software – including Prometheus, Mimir, Loki, and Tempo – without the overhead of installing, maintaining and scaling their own observability stack.
The Databases team owns the telemetry databases that are Mimir for metrics, Loki for logs, Tempo for traces, and Pyroscope for profiles. Our databases are OSS projects that we also offer as a Cloud service. They are multi-tenant distributed systems implemented in Go and running on Kubernetes across all major Cloud service providers.
What You’ll Be Doing
- Take an active role in influencing our roadmap and your own career objectives.
- Work with your team to deliver new features, then use the results to iterate and improve.
- Drive projects from initial ideation all the way to operations once it is in the hands of customers.
- Embrace our open-source culture and contribute to other projects that may not directly fall within your team’s scope.
- Design, build, operate, and maintain critical systems, owning the reliability, performance, and availability.
- Participate in on-call rotations and take ownership of the services you’re running.
- Support other team members, participate in design discussions and collaborate with the team.
- Learn new skills by gaining a deeper understanding of our cloud product and our customers.
Requirements
- Solid experience with at least one programming language (Go is our primary language, but familiarity with Python, C, C++, or Rust is applicable).
- Experience with delivering projects from gathering requirements and brainstorming ideas to shipping products to customers in a self-driven way.
- Experience with developing software that runs in the Cloud or systems engineering experience.
- Experience writing clean, robust, and performant software that is easily maintained by others.
Bonus Points For
- Experience working with Kubernetes and/or Docker.
- Prior use of Grafana and Prometheus in operational roles.
- Exposure to microservices architecture and distributed systems.
- Familiarity with being on-call, SRE tasks, or infrastructure as code.