Job SummaryWe are seeking a Site Reliability Engineer to support Cyber Data Risk & Resilience by ensuring the reliability, availability, performance, and operational visibility of critical cybersecurity platforms and services. This role is responsible for maintaining production systems, building observability solutions, automating operational processes, supporting incident response, enhancing executive dashboards, and driving continuous improvements across cloud-based and distributed environments.
Key Responsibilities- Maintain and improve the reliability, availability, scalability, and performance of cybersecurity platforms and supporting infrastructure.
- Monitor system health, identify operational risks, respond to incidents, and drive timely resolution of service-impacting issues.
- Instrument infrastructure, applications, APIs, databases, cloud components, and data pipelines for end-to-end observability.
- Design, build, and enhance monitoring, alerting, logging, tracing, and dashboard solutions across distributed systems.
- Develop actionable alerts that reduce noise, improve signal quality, and accelerate incident response.
- Define and monitor SLIs, SLOs, SLAs, error budgets, latency, throughput, availability, and operational risk metrics.
- Build, maintain, and improve operational dashboards for engineering, operations, cybersecurity, risk, and executive leadership.
- Continuously enhance executive dashboards with service health, reliability trends, incident reporting, and operational performance metrics.
- Partner with engineering, cloud, infrastructure, application, and cybersecurity teams to improve platform reliability.
- Participate in incident response, root cause analysis, post-incident reviews, and problem management activities.
- Automate operational tasks, health checks, reporting, deployment validation, and recovery procedures.
- Support CI/CD pipelines, DevOps processes, release readiness, rollback validation, and production support activities.
- Contribute to resiliency engineering initiatives including capacity planning, performance tuning, disaster recovery, failover testing, and resilience validation.
- Ensure monitoring, dashboards, alerting, and operational processes comply with enterprise security, governance, and risk standards.
- Perform other duties as assigned.
Required Qualifications- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent work experience.
- 10+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Systems Engineering, Software Engineering, or Production Operations.
- Experience supporting highly available, distributed, cloud-based, or mission-critical technology platforms.
- Strong experience with monitoring, observability, logging, tracing, dashboards, and service health reporting.
- Experience instrumenting applications, infrastructure, cloud services, APIs, databases, and distributed systems.
- Strong understanding of SRE concepts including SLIs, SLOs, SLAs, error budgets, capacity planning, and incident management.
- Experience designing meaningful monitoring and actionable alerting strategies.
- Strong scripting or programming skills using Python, Java, Bash, PowerShell, or similar languages.
- Experience with cloud platforms including AWS, Azure, or GCP.
- Experience with Infrastructure as Code tools such as Terraform.
- Experience supporting CI/CD pipelines, DevOps workflows, release management, and production environments.
- Experience troubleshooting distributed systems, REST APIs, event-driven architectures, messaging platforms, and service integrations.
- Familiarity with relational and NoSQL databases such as PostgreSQL, Microsoft SQL Server, MongoDB, or similar technologies.
- Strong analytical, troubleshooting, problem-solving, and communication skills.
Preferred Qualifications- Experience supporting cybersecurity, risk, resilience, or enterprise security platforms.
- Experience with observability platforms such as Splunk, Grafana, Prometheus, Datadog, Dynatrace, New Relic, Azure Monitor, CloudWatch, or OpenTelemetry.
- Experience building executive-level dashboards, operational scorecards, and reliability reporting.
- Experience with automated health checks, synthetic monitoring, dependency mapping, and operational runbooks.
- Experience with Kubernetes, Docker, serverless platforms, or cloud-native technologies.
- Experience with Apache Kafka or other messaging platforms.
- Familiarity with cloud security governance tools such as Azure Policy, AWS SCP, Wiz, Prisma, CloudGuard, or similar solutions.
- Experience with AI cloud platforms such as Azure AI, AWS Bedrock, or Google Vertex AI.
- Experience supporting Linux and Windows environments through scripting and automation.
Required Skills- Site Reliability Engineering (SRE)
- DevOps
- Cloud Platforms (AWS, Azure, GCP)
- Infrastructure as Code (Terraform)
- Monitoring & Observability
- Logging & Tracing
- Alerting
- Dashboard Development
- Incident Management
- Root Cause Analysis
- Service Reliability
- SLIs / SLOs / SLAs
- Performance Monitoring
- Capacity Planning
- Disaster Recovery
- CI/CD
- Python
- Java
- Bash
- PowerShell
- REST APIs
- Distributed Systems
- Microservices
- SQL Databases
- NoSQL Databases
- Automation
- Troubleshooting
- Cloud Infrastructure
- Operational Excellence
- Communication Skills
- Problem Solving
Preferred Skills- Splunk
- Grafana
- Prometheus
- Datadog
- Dynatrace
- New Relic
- Azure Monitor
- AWS CloudWatch
- OpenTelemetry
- Kubernetes
- Docker
- Apache Kafka
- Wiz
- Prisma Cloud
- CloudGuard
- Azure AI
- AWS Bedrock
- Google Vertex AI
- Linux Administration
- Windows Administration
- Executive Dashboard Reporting
- Synthetic Monitoring
- Service Dependency Mapping
- Cloud Security
- Cybersecurity Platforms
EducationBachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent work experience.
CertificationsRelevant cloud, DevOps, SRE, Kubernetes, or cybersecurity certifications are preferred.