Overview/ Job ResponsibilitiesWe are seeking an experienced Senior Site-Reliability Engineer to join our infrastructure team. The ideal candidate will be responsible for ensuring the reliability, availability, and performance of our Windows-based production environments. You will bridge development and operations to deliver highly available services while maintaining operational excellence.
Key Responsibilities:
- Design, implement, and maintain scalable infrastructure using Infrastructure as Code (IaC) practices
- Develop and maintain automation scripts using various scripting languages (PowerShell, Python, Ruby, etc) for OS provisioning, configuration management, and operational tasks
- Implement and manage configuration management solutions (Terraform, Puppet, and/or Chef) across hybrid environments
- Monitor system health, performance, and availability using industry-standard tools and practices
- Establish and enforce SLAs, SLOs, and error budgets for production services
- Participate in on-call rotation and respond to incidents with a focus on rapid restoration and root cause analysis
- Collaborate with development teams to improve deployment pipelines and release processes
- Document operational procedures, runbooks, and architectural decisions
- Conduct post-mortem reviews and implement corrective actions to prevent recurrence
Minimum Qualifications- 5+ years of experience in Systems Administration, DevOps, or Site-Reliability Engineering roles
- Strong expertise in Windows Server environments (2016+), including Active Directory, IIS, and MS SQL.
- Strong cloud skills, AWS experience preferred.
- Advanced proficiency in scripting, including module development and integration with REST APIs
- Hands-on experience with configuration management tools:
- Terraform (infrastructure provisioning)
- Puppet or Chef (configuration management)
- Experience with monitoring and observability platforms (e.g., Prometheus, Grafana, Datadog, New Relic)
- Solid understanding of networking concepts (DNS, TCP/IP, Load Balancing, VPN)
- Strong problem-solving skills with the ability to troubleshoot complex issues across multiple technology layers
Desired Qualifications- Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent professional experience)
- Certifications such as AWS Solutions Architect, Microsoft Certifications, or HashiCorp Certified: Terraform Associate
- Experience with containerization technologies (Docker, Kubernetes)
- Familiarity with CI/CD tools (Gitlab Pipelines, Jenkins, GitHub Actions)
- Knowledge of security best practices and compliance frameworks
- Experience with log aggregation and analysis tools (ELK Stack, Splunk)
The estimated annual salary range for this position is $105,000.00 - $140,000.00/year. Actual compensation will be determined based on factors such as experience, qualifications, skills, geographic location, contract requirements, and business needs.
Entarian offers a comprehensive benefits package to eligible employees, including medical, dental, and vision insurance; life, AD&D, and disability insurance; paid time off and 11 company holidays; 401(k) retirement plan with company matching; and additional employee benefits and wellness resources. Benefits and compensation are subject to applicable eligibility requirements and the terms of Entarian's plans and policies.