Senior Site Reliability Engineer - Observability Engineer | NordVPN

Nord Security

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years experience in distributed systems observability
  • Strong knowledge of monitoring architecture and signal design
  • Ability to design alerts that minimize noise and enhance engineer trust
  • Proficient in Python for scripting and automation
  • Skilled in Linux administration and debugging
  • Understanding of networking fundamentals

Responsibilities

  • Design and enhance monitoring pipelines across global infrastructure
  • Implement service-level monitoring based on key performance indicators
  • Develop actionable alerts to reduce alert fatigue
  • Maintain custom exporters and scripts for metrics collection
  • Collaborate with data team for anomaly detection and insights
  • Decipher service signals for impactful measurement

Benefits

  • Work alongside global cybersecurity experts impacting millions
  • Access extensive training programs and mentorship opportunities
  • Flexible hybrid work model with option to work from anywhere
  • Physical wellness programs and access to fitness resources
  • Mental health support including free consultations and wellness apps
  • Private health insurance for peace of mind
  • Celebrate personal milestones with special gifts
  • Engage in team-building activities and company events
  • Participate in a memorable company getaway with fun activities
Full Job Description
NordVPN runs a global edge infrastructure serving millions of users. Knowing what's happening across that infrastructure - in real time, at scale, without drowning in noise - is what this role exists to solve. We're looking for a Senior Site Reliability Engineer (SRE) focused on observability: designing monitoring systems, improving signal quality, reducing alert fatigue, and collaborating with data teams on anomaly detection. You'll own how we understand the health and behavior of our distributed systems. **Main responsibilities** - Design, build, and improve monitoring pipelines and observability tooling across globally distributed infrastructure - Define and implement service-level monitoring based on golden signals (latency, traffic, errors, saturation) - Reduce alert fatigue - build meaningful, actionable alerts that engineers trust - Develop and maintain custom exporters, scripts, and integrations for metrics and log collection - Collaborate with the data team on anomaly detection and data-driven operational insights - Understand service signals - know what to measure, why, and what the numbers actually mean **Core requirements** - Distributed systems observability - monitoring architecture, signal design, dashboarding - Golden signal thinking - you design monitoring around what matters, not what's easy to measure - Alert design - reducing noise, building actionable alerts, managing on-call sanity - Python - scripting, custom exporters, automation, data processing - Linux administration and debugging - Networking fundamentals **Bonus Points For** - SaltStack - Advanced networking - traffic analysis, protocol-level debugging - Advanced data knowledge - aggregation strategies, downsampling, cardinality management, retention trade-offs - Proven track record of onboarding new systems/services into monitoring from scratch - Familiarity with agentic engineering - Claude Code, LLM integrations, MCP workflows **Tools You Will Use** - Naemon (Nagios) and Gearmand - Prometheus-based exporters - Telegraf - Fluent Bit - VictoriaMetrics ecosystem - OpenSearch - Grafana **What We Offer** **Innovate with industry leaders** Work alongside global experts to build world-leading cybersecurity tools, impacting millions of users around the world. **Learn & grow** Boost your skills via our extensive training programs (online and offline) & other resources. Benefit from mentorship and career-switch opportunities to grow within the company. **Hybrid work** Enjoy the flexibility with 3 office days and working from home for the remaining 2. **Work from anywhere** Recharge with a change of scenery - choose work from any location when you feel a need to power your creativity and drive. **Physical well-being** Fuel your active lifestyle with online workouts led by our Physical Well-Being experts. Unlock a variety of sports and wellness facilities, like gyms, swimming pools, and fitness classes, with the Multisport card. **Mental & emotional health** Nurture your mind with free psychologist consultations, dedicated mental health events, and premium access to top-rated wellness apps like Calm, Headspace, and Mindletic. **Premium healthcare** Receive private health insurance giving you peace of mind for your health needs. **Joyful moments - special treats** Celebrate life's big moments with special gifts from us on your birthday, anniversary, and other major events, such as weddings or the arrival of a new family member. **Company events & team-building** Experience iconic Nord Security celebrations, team-buildings, and knowledge-sharing events, nurturing bonds that fuel our success. **Workation** Embark on a legendary company getaway abroad, filled with exciting activities, live concerts, engaging workshops, and epic time together.

Similar Jobs

More Jobs at Nord Security

More Information Technology Jobs

Find similar Senior Site Reliability Engineer - Observability Engineer | NordVPN jobs: