Sr. Site Reliability Engineer

Blackpoint Cyber

• $135K — $160K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience as a Senior Site Reliability Engineer or similar role, emphasizing cloud infrastructure management and automation.
  • Expertise in Infrastructure as Code using Terraform and Terragrunt for enterprise deployments.
  • In-depth knowledge of AWS, focusing on secure and scalable architecture management.
  • Hands-on experience with distributed data streaming technologies like Confluent Cloud and Apache Kafka.
  • Experience with Redis for caching and Amazon RDS for database management.
  • Familiarity with analytics platforms like OpenSearch or Elasticsearch and proficiency in monitoring infrastructure with tools like Prometheus and Grafana.
  • Strong problem-solving skills and experience in Agile environments.

Responsibilities

  • Design, develop, and maintain scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt).
  • Optimize AWS environments for cost efficiency, security, and availability.
  • Administer Kubernetes clusters and support continuous delivery through tools like Helm and ArgoCD.
  • Scale and manage data streaming infrastructure for enterprise-level processing.
  • Deploy and maintain Redis for real-time data handling and caching.
  • Implement monitoring and incident response frameworks utilizing Prometheus and Grafana.
  • Collaborate with development teams for smooth integration of services and troubleshooting.

Benefits

  • Health, Vision, and Dental Insurance plans, plus Life Insurance options for eligible employees in the U.S.
  • Robust 401k plan to support employee financial growth.
  • Discretionary Time Off for better work-life balance.
  • Equity participation available to employees globally, sharing in the company's success.
Full Job Description
SUMMARY

We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance. You'll work across cloud platform administration, container orchestration, data streaming, observability, and incident response - partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster a culture of continuous improvement.

RESPONSIBILITIES
  • Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.
  • Own and optimize our AWS cloud environment, ensuring cost efficiency, security best practices, and high-availability standards.
  • Manage and optimize Kubernetes cluster environments (Helm, ArgoCD, Istio, Kustomize) to support continuous delivery and infrastructure-as-code practices.
  • Administer and scale data streaming infrastructure (Confluent Cloud, Apache Kafka) to support enterprise-level data processing.
  • Deploy, configure, and maintain Redis for caching and real-time data processing.
  • Implement and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, Alert Manager, Grafana CloudOpsGenie/PagerDuty) to ensure system reliability and performance.
  • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.
  • Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure.
  • Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance and maximize uptime.
  • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies to enhance scalability, reliability, and maintainability.
  • Stay current on emerging SRE trends and tools, andtools and help the team adopt relevant industry advancements and best practices.

REQUIREMENTS
  • 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation.
  • Expertise in Infrastructure as Code (Terraform, Terragrunt) for enterprise-scale deployments.
  • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures.
  • Extensive hands-on experience with distributed data streaming (Confluent Cloud, Apache Kafka).
  • Proven experience with Redis for caching and Amazon RDS for relational database management.
  • Experience with enterprise search and analytics platforms (OpenSearch, Elasticsearch, ChaosSearch).
  • Proficiency designing and implementing monitoring/alerting infrastructure (Prometheus, Grafana, Alert Manager, Grafana Cloud, OpsGenie/PagerDuty).
  • Practical experience with feature flag systems (LaunchDarkly/PostHog) for controlled release management.
  • Extensive experience administering production-grade Kubernetes (Helm, ArgoCD, Istio); working knowledge of Kustomize.
  • Strong problem-solving skills, with the ability to troubleshoot complex systems in production.
  • Strong communication and collaboration skills, with experience working in Agile environments.

NICE TO HAVE
  • Experience with Terragrunt to manage Terraform across multiple environments.
  • Extensive hands-on experience with distributed data streaming (Kafka).
  • Multi-cloud experience (Google Cloud Platform, Microsoft Azure).
  • Understanding of security frameworks and compliance standards for cloud-native/containerized environments.
  • Serverless computing and CI/CD pipeline experience (Jenkins, GitHub Actions).
  • Software development proficiency in Node.js, Python, and/or Go.

For eligible employees in the US, Blackpoint offers competitive Health, Vision, Dental, and Life Insurance plans, a robust 401k plan, Discretionary Time Off, and other minor perks. International employees receive competitive benefits in accordance with local market standards and applicable country requirements.

Blackpoint believes all employees should share in the company's success - equity participation is available to employees globally, with program details varying by location and employment structure.

Similar Jobs

More Jobs at Blackpoint Cyber

More Information Technology Jobs

Find similar Sr. Site Reliability Engineer jobs: