Senior Cloud Engineer - Observability Platform

Compunnel

$120K — $150K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of hands-on experience with distributed multi-tiered systems
  • Deep understanding of telemetry architecture
  • Proficient in observability tools like Open Telemetry
  • Strong programming skills in Python, Java, or Go
  • Experience with container orchestration in Kubernetes
  • Expertise in infrastructure as code (Terraform, CFT)
  • Solid grasp of networking concepts and cloud deployment strategies

Responsibilities

  • Deploy and manage enterprise time series systems like Prometheus
  • Create and implement best practices with PromQL and Grafana
  • Troubleshoot application performance issues autonomously
  • Develop features for performance and reliability optimizations
  • Automate the creation of intelligent monitors and SLOs
  • Support strategic enhancements for platform security and resiliency
  • Collaborate within a diverse and innovative team environment

Benefits

  • Dedicated learning day each week
  • Empowered and collaborative team culture
  • Focus on open-source contributions
  • Continuous learning and innovation environment
  • Opportunities to influence cloud technology advancements
Full Job Description
JOB SUMMARY Do you want to work on leading edge cloud technologies which are transforming how developers work with cloud and improve efficiency of software through enabling AI features for Observability products? As a Senior Cloud Engineer, you will work within a diverse team comprised of passionate technologists who believe in the power of innovation and constant collaboration. We believe that small, empowered, self-motivated teams can achieve outstanding things. We are passionate about opensource contributions, sharing our expertise and knowledge with the engineering community while adopting a continuous learning approach supported by a dedicated learning day each week. Working on Foundation & Pipelines team, which supports the Core & Common initiative of building, deploying, and supporting our new OTel observability platform focused on metrics and traces. This role works toward to strategic platform enhancements, including platform security and resiliency and new features. Key Responsibilities In this role, you will be responsible for deploying, configuring, and managing enterprise time series system like Prometheus, creating best practices using PromQL and Grafana. Debug and investigate application performance issues down to the root cause, as both a developer assistant and a fully autonomous agent. Proactively build features which recommend performance and reliability-based optimizations to prevent the next incident. Automatically create intelligent monitors and SLOs for the most important business flows and critical paths. Required Qualifications 5+ years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale. Deep understanding of distributed systems and telemetry architecture. Experience choosing/modeling the right technique for the job (e.g., anomaly detection, ranking/recommendation, NLP), and knowing when a heuristic beats a model. Experience operating observability pipelines in Kubernetes or similar orchestration environments. Hands-on experience in 2 or more languages (Python, Java, Go etc.). Hands-on experience with Open Telemetry (OTEL) or any opensource Observability implementations. 5+ years of experience leading projects and designing, analyzing, and troubleshooting distributed systems. Create and maintain Grafana dashboards, visualizations, and alerts for real-time operational insights. Hands-on experience designing and building scalable and resilient applications in the cloud. Good knowledge on networking concepts such as DNS, Load Balancers, routers , Linux etc. Experience on containerization technologies such as Docker, Kubernetes etc. Experience in deploying applications in AWS platforms such as EC2 & EKS or equivalent platforms from any other cloud provider such as GCP, Azure etc. Experience delivering software using engineering best practices and principles. Extensive knowledge of infrastructure as code (Terraform, CFT, CDK, etc.). Hands-on experience with continuous integration and continuous delivery/deployment using ALM tools such as Jenkins. Preferred Qualifications Proven experience delivering LLM/agent features to production (prompting, tooling, evals, safety/guardrails). Experience working on Logging, Metrics, and Tracing frameworks and understanding of Logging and Metrics Data Models. You have proven ability to use AI coding tools in day-to-day workflows and validate, critique, and refine AI-generated output. Experience operating and/or using modern observability tools at scale (Prometheus, Grafana, ELK, Jaeger, Open Telemetry, fluentD, fluentBit, Open Tracing, Datadog, Splunk, etc...) Certifications

Similar Jobs

More Jobs at Compunnel

More Information Technology Jobs

Find similar Senior Cloud Engineer - Observability Platform jobs: