Role Overview:Seeking a skilled Grafana Engineer to design, implement, and maintain enterprise monitoring and observability solutions using Grafana and related technologies. The ideal candidate will have hands-on experience in building dashboards, configuring alerts, integrating multiple data sources, and supporting cloud-native environments. The role involves collaborating with DevOps, SRE, Infrastructure, and Application teams to ensure high availability, performance, and reliability of business-critical applications and platforms.
Key Responsibilities:- Design, develop, and maintain Grafana dashboards, visualizations, and reports for infrastructure, applications, and business metrics.
- Configure and manage Grafana Alerting for proactive monitoring and incident response.
- Integrate Grafana with various data sources such as: Prometheus, Loki, Elasticsearch/OpenSearch, InfluxDB, SQL databases, Cloud monitoring platforms (AWS CloudWatch, Azure Monitor, Google Cloud Operations).
- Develop observability solutions covering metrics, logs, and traces.
- Implement monitoring strategies for Kubernetes, containers, virtual machines, cloud infrastructure, and enterprise applications.
- Collaborate with SRE, DevOps, Application Support, and Platform Engineering teams to define monitoring requirements.
- Automate dashboard deployment and configuration using Infrastructure as Code (IaC) tools.
- Tune monitoring systems to minimize alert fatigue and improve operational efficiency.
- Perform root cause analysis using monitoring and logging data.
- Create and maintain technical documentation, monitoring standards, and operational runbooks.
- Support capacity planning, performance analysis, and system optimization initiatives.
- Ensure security and governance for monitoring platforms and data access.
Required Skills:- Strong experience with Grafana dashboard development and administration.
- Expertise in Grafana Alerting, notification channels, and alert rule management.
- Experience with Grafana Enterprise.
- Experience with: Prometheus, Loki, Tempo, OpenTelemetry, Elasticsearch/OpenSearch, Splunk (preferred).
- Understanding of Metrics, Logs, and Distributed Tracing concepts.
- Experience with one or more cloud platforms: AWS, Azure and GCP.
- Familiarity with Kubernetes and container orchestration.
- Knowledge of CI/CD pipelines and DevOps practices.
- Proficiency in scripting languages such as: Python, Bash/Shell, PowerShell.
- Experience with Terraform, Ansible, or similar automation tools.
- Experience querying and analyzing monitoring data.
- Strong analytical and troubleshooting skills.
- Excellent communication and stakeholder management.
- Ability to work independently and collaboratively.
- Problem-solving mindset with attention to detail.
- Strong documentation and knowledge-sharing capabilities.
- Grafana Enterprise deployment experience.
- OpenTelemetry implementation experience.
- Experience with AIOps and observability platforms.
- Exposure to application performance monitoring (APM) tools such as Dynatrace, AppDynamics, Datadog, or New Relic.
Qualifications: