Job Summary:
The Site Reliability Engineer will provide enterprise-level systems administration, monitoring, automation, deployment, and production support for complex distributed environments. The role requires strong experience building automation scripts, proactive monitoring dashboards, alerting solutions, SDLC practices, and high-availability architectures. The ideal candidate will be a seasoned SRE professional with strong troubleshooting skills across Windows and Linux environments, networking, databases, and enterprise applications, combined with a strong sense of ownership and the ability to independently resolve complex technical issues.
Key Responsibilities
• Perform enterprise systems administration, monitoring, troubleshooting, tuning, and deployment activities.
• Develop automation scripts and process improvements to increase operational efficiency and reliability.
• Build application dashboards and configure proactive monitoring and alerting using tools such as Grafana, Influx, Datadog, Moog, and ThousandEyes.
• Establish monitoring and alerting capabilities to identify potential issues early and improve system reliability.
• Support and improve Software Development Lifecycle (SDLC) practices and operational processes.
• Administer and troubleshoot Windows Server 2016, 2019, and 2022 environments hosted on virtual machines.
• Perform Linux and Windows system administration, troubleshooting, and performance tuning.
• Troubleshoot enterprise infrastructure and application issues across distributed environments.
• Support large-scale distributed systems and high-availability architectures.
• Develop automation and operational tooling using programming or scripting languages such as .NET, C#, PowerShell, YAML, Java, Python, or Bash.
• Support and troubleshoot enterprise databases, including SQL, Oracle, MongoDB, and PostgreSQL environments.
• Apply IP networking knowledge to troubleshoot DNS, DHCP, firewalls, IP routing, and related connectivity issues.
• Collaborate with technical teams to proactively identify risks, resolve incidents, and improve system reliability.
• Take ownership of complex technical issues and drive them through resolution in distributed environments.
Required Qualifications
• 6-8+ years of experience in enterprise-level systems administration and support.
• 6-8+ years of experience developing automation scripts and building application dashboards and proactive monitoring solutions.
• 6-8 years of experience practicing SDLC processes and driving process improvements.
• Hands-on experience with enterprise systems administration, monitoring, and deployment activities.
• Strong experience with Windows Server 2016, 2019, and 2022 hosted via virtual machines.
• Strong knowledge of IP networking concepts, including DNS, DHCP, firewalls, and IP routing.
• Familiarity with large-scale distributed systems and high-availability architectures.
• Strong Linux and Windows system administration, troubleshooting, and performance-tuning experience.
• Development or scripting experience with one or more languages or technologies such as .NET, C#, PowerShell, YAML, Java, Python, or Bash.
• Knowledge of one or more database technologies, including SQL, Oracle, MongoDB, or PostgreSQL.
• Bachelor's degree in Computer Science or a related discipline.
• Strong customer orientation with the ability to proactively own, communicate, and follow through on projects and issues.
• Strong sense of ownership and accountability for resolving problems in distributed environments.
• Ability to investigate complex technical issues deeply and systematically.
• Self-starter with the ability and confidence to independently resolve issues and communicate results to the broader team.
Preferred Qualifications
• Experience working in the financial services industry.
• Experience working with Agile methodologies.
• Experience supporting complex trading or similarly high-availability technology environments.