Job SummaryThe Role Senior DNS Engineer - SRE is an expert level technical role responsible for the strategic architecture, global scalability, and ironclad reliability of the mission-critical DNS infrastructure powering our ISP and core network services. This position is designed for a veteran engineer who masters the Site Reliability Engineering (SRE) philosophy-championing automation, observability, and self-healing systems as the standard for Tier-1 operations. You will mentor junior staff and lead the technical roadmap, combining deep IP networking expertise with modern security protocols to defend against sophisticated threats while maintaining sub-millisecond performance for millions of users.
The Impact This is a high-visibility, influential role. As a technical authority, you will drive cross-functional initiatives across Product, Security, and Service Assurance teams. Your mission is to evolve a carrier-grade DNS ecosystem that sets the industry standard for privacy (DoH/DoT) and availability. You will not only manage infrastructure but also define the engineering culture, ensuring our global name services remain resilient, compliant, and ahead of the technological curve.
Responsibilities- Core Platform Strategy, Architectural Vision: Lead the high-level design and evolution and operation of global, multi-site DNS architectures. Implement advanced Anycast routing strategies, multi-provider redundancy, and zero-touch failover mechanisms.
- Architectural Ownership: Lead the design and evolution of global DNS architectures, ensuring high availability through Anycast routing, multi-provider redundancy, and automated failover mechanisms.
- Strategic Vendor Relations: Act as the primary technical authority in engagements with DNS and infrastructure vendors, driving roadmaps that align with our long-term reliability and security goals.
- Lifecycle & Capacity Management: Oversee the full lifecycle of DNS platforms-including automated software deployments, hardware refreshes, and proactive capacity planning-to stay ahead of traffic growth.
- Standardization & Policy: Optimize, Define and enforce organization-wide standards for DNS record management, security protocols (DNSSEC), and traffic steering policies to optimize user latency.
- Reliability Engineering: Convert "Strategic Design" into "Operational Reality" by defining Service Level Objectives (SLOs) and Error Budgets for all core name services.
Advanced DNS Operations & SRE
- Protocol Mastery: Serve as the final escalation point for complex protocol issues involving recursion, iteration, and specialized record (A, AAAA, CNAME, MX, TXT, SRV) in a dual-stack environment.
- Security Engineering: Architect robust DNSSEC implementations and global DDoS mitigation strategies. Lead the response to DNS-based threats, including cache poisoning and sophisticated amplification attacks.
- Toil Elimination & Tooling: Lead the development of custom automation frameworks using industry tools. Replace all manual intervention with robust, tested workflows via Ansible, Terraform, or Kubernetes operators.
- Full-Stack Performance Tuning: Optimize Linux kernel parameters for extreme network throughput. Conduct forensic-level analysis of BIND, Unbound, or PowerDNS logs and telemetry to identify latent systemic issues.
- Deep Observability: Design comprehensive monitoring ecosystems using Prometheus, Grafana, and dnstap. Transform raw query data into executive-level insights regarding latency, error rates (NXDOMAIN/SERVFAIL), and user behavior.
Qualifications- Education: Bachelor's or Master's degree in Computer Science, Telecommunications, or a related field (or equivalent practical experience)
- Experience: 5+ years in networking or systems engineering, with at least 3 years focused on SRE principles within a high-scale production environment
- DNS Authority: Expert-level experience configuring and maintaining at least three of the following: BIND, Unbound, PowerDNS, AWS Route 53, or Azure DNS
- Advanced Networking: Expert understanding of TCP/IP, BGP (for Anycast), and encrypted transport protocols (DoH/DoT)
- Software Engineering: High proficiency in Python, Go, or Bash, with a proven track record of building production-grade automation tools
- Observability Stack: Deep experience architecting monitoring solutions with Prometheus, Grafana, and OpenTelemetry
Preferred Qualifications
- ISP/Carrier Experience: Previous experience in a Tier-1 or Tier-2 service provider environment managing infrastructure for millions of subscribers
- Cloud-Native DNS: Proven success in managing hybrid-cloud DNS environments and migrating legacy workloads to cloud-native or containerized solutions
- Security Certifications: Relevant certifications (e.g., CISSP, CCNP/CCIE, or SRE-specific certifications) are a plus
- Leadership: Experience mentoring junior engineers and leading complex, multi-quarter technical projects from inception to completion
Pay is competitive and based on a number of job-related factors, including skills and experience. The starting pay rate/range at time of hire for this position in New York is 100,246.00 - 164,689.00 / year. For other locations, please inquire with your recruiter. The rates/ranges provided herein are the anticipated pay at the time of hire, and do not reflect future job opportunity.