Site Reliability Engineer II
Day in the Life of the Site Reliability Engineer II:
The Site Reliability Engineer IIoptimizesservice performance, activelyparticipatesin reliability improvements, and conducts in-depth SLO and capacity analysis. This position exists to enhance system reliability and scalability while contributing to automation and self-service tool development..
- Performance Monitoring: Monitor service performance,assistin troubleshooting production issues, and learn system architecture.
- Reliability Participation: Monitor service reliability,participatein resolving basic issues, and learn disaster recovery testing procedures.
- SLO Implementation: Understand SLO concepts,monitorand analyze SLO patterns, andassistin implementing SLO visualization and alerting.
- Capacity Analysis: Perform basic capacity analysis,identifytrends in system capacity, andparticipatein capacity planning.
- Automation Deployment: Deploy andmaintainexisting automation tools, create simple scripts, and troubleshoot automation scripts.
Required Qualifications - About you:
We are looking for candidates who possess the combination of the following achievements, skills, and behaviors:
- 5+ years of experience in enterprise networking, including handson work with routing, switching, firewalls, load balancers, and VPN technologies.
- Strong understanding of cloud networking architectures across including VPC/VNet design, peering, private link, and hybrid connectivity models.
- Experience with network security technologies, such as security groups, NACLs, firewall policies, WAF, IDS/IPS, and microsegmentation.
- Proficiency in Layer 2 and Layer 3 network protocols, including BGP, OSPF, EIGRP, DNS, DHCP, NAT, and IP addressing/subnetting.
- Handson experience with load balancers and ingress technologies, including F5, NGINX, Azure Application Gateway, ALB/NLB, or equivalent.
- Strong troubleshooting skills using packet analyzers tools, flow logs, and network monitoring platforms.
- Skilled in analyzing performance trends andidentifiesoptimization opportunities.
- Collaborates with teams to improve monitoring coverage.
- Ability toparticipatein structured reliability testing and analysis.
- Able to evaluate system components for resilience.
- Contributes to reliability-focused design discussions.
- Skilled in analyzing trends to inform service improvements.
- Collaborates with teams to align SLOs with user expectations.
- Develops moderately complex automation tools.
- Skill in building internal self-service capabilities.
- Evaluates automation opportunities for operational efficiency.
- Skilled in analyzing capacity data to inform scaling decisions.
- Able to recommend improvements for resourceutilization.
- Ensures scalability is considered in feature development.
- Follow predefined procedures to deploy PROS products and third-party applications to the Cloud environments.
- Contribute to the release management documentation.
- Gain understanding of application architecture and interaction between system components.
Highly Preferred:
- Bachelors Degree in Computer Science, Information Technology, or a related field
- Practical experience with Fortigate firewalls and F5 appliances is highly desirable
AI Fluency & Growth Mindset- We welcome candidates who:
- Understand core AI concepts and apply them ethically to enhance productivity, insights, and decision-making.
- Craft effective prompts to optimize the quality and relevance of AI-generated outputs.
- Explore and apply agentic AI systems, using or managing autonomous agents to streamline workflows and automate tasks.
- Leverage AI tools to boost efficiency, creativity, and innovation in their daily work.
- Stay curious and adaptable, continuously experimenting with AI-driven solutions to elevate team performance and customer impact.