Nextpoint

Site Reliability Engineer

Nextpoint$110K — $130K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, preferably in B2B SaaS
  • Hands-on experience with AWS services such as EC2, S3, RDS, and Lambda
  • Knowledge of infrastructure-as-code tools like Terraform or CloudFormation
  • Skills in scripting or programming (Python or Go) for automation
  • Experience with containerization technologies (Docker, Kubernetes)
  • Operational experience with CI/CD pipelines
  • Familiarity with monitoring tools (Datadog, Prometheus) and security compliance frameworks (SOC 2, HIPAA)

Responsibilities

  • Maintain and extend infrastructure as code for existing systems
  • Support and enhance CI/CD pipelines
  • Monitor cloud cost trends and identify optimization opportunities
  • Monitor performance metrics for production systems
  • Respond to production incidents during US business hours
  • Triage and resolve infrastructure requests and incidents
  • Communicate technical issues clearly to technical and non-technical stakeholders

Benefits

  • Opportunity to work with cutting-edge AI technologies
  • Hands-on role with significant impact on platform reliability and customer trust
  • Collaborative environment partnering with engineering teams on new features
  • Engagement in compliance and security activities ensuring data integrity
  • Exposure to large-scale legal data processing
Full Job Description
ABOUT THE ROLE

Nextpoint's platform handles massive, unpredictable volumes of sensitive legal data - unlimited-upload document review, AI-assisted analysis, and secure electronic production - all running on AWS with zero downtime tolerance for firms in active litigation. We're looking for a Site Reliability Engineer to help maintain the reliability, scalability, and security posture of that platform as we expand our AI capabilities (built on Amazon Bedrock) and grow our customer base.

This is a hands-on role for someone who wants to be the first line of response for production infrastructure at a company where reliability is a customer-trust issue, not just an engineering metric.

RESPONSIBILITIES

Infrastructure & Automation (50%)
  • Maintain and extend existing infrastructure-as-code (Terraform/CloudFormation/CDK) following established patterns and standards
  • Support and operate CI/CD pipelines; implement improvements as directed
  • Monitor cloud cost trends and flag optimization opportunities for review


Reliability & Operations (25%)
  • Monitor uptime, latency, and performance SLOs/SLIs for production systems supporting document upload, processing, review, and production workflows
  • Participate in the on-call rotation and serve as first responder for production incidents during US business hours
  • Triage, troubleshoot, and resolve incoming infrastructure requests and incidents; escalate and coordinate on complex root-cause work
  • Write clear post-incident reports and help investigate recurring incidents and cost overruns
  • Maintain and extend existing monitoring, alerting, and observability tooling

Cross-Functional Collaboration (15%)
  • Partner with engineering teams to build reliability, scalability, and observability into new features from design through launch
  • Document runbooks, architecture decisions, and operational procedures for the broader engineering team
  • Communicate incident status and technical issues clearly to engineering and non-technical stakeholders
  • Participate in design reviews to flag reliability or operational concerns early

Security & Compliance (10%)
  • Support SOC 2 compliance activities and help uphold encryption, access-control, and audit-trail standards across all environments
  • Implement and maintain security best practices for infrastructure handling confidential legal and client data, including AI workloads on Amazon Bedrock
  • Support security reviews, vulnerability management, and patching cadences across production systems
  • Maintain permissions-based access controls and comprehensive audit logging in line with client security commitments


QUALIFICATIONS
  • 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles, ideally in a B2B SaaS environment
  • Deep hands-on experience with AWS (EC2, S3, RDS, Lambda, VPC, IAM, CloudWatch, or equivalent services)
  • Working knowledge of infrastructure-as-code (Terraform, CloudFormation, or CDK) and configuration management
  • Proficiency in at least one scripting/programming language (Python, Go, or similar) for automation and tooling
  • Experience with containerization and orchestration (Docker, Kubernetes, or ECS)
  • Track record of operating CI/CD pipelines
  • Experience with monitoring/observability stacks (Datadog, CloudWatch, Prometheus/Grafana, or similar)
  • Familiarity with security compliance frameworks (SOC 2, HIPAA, or similar) and encryption/access-control best practices
  • Experience supporting systems handling large-scale, variable-volume data processing is a plus
  • Exposure to AI/ML infrastructure (e.g., Amazon Bedrock, model-serving pipelines) is a plus, given our growing AI feature set
  • Experience using ClaudeCode, Kiro, OpenCode or similar Agentic AI
  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)
  • Strong written and verbal communication skills, with the ability to write clear runbooks and explain technical tradeoffs to non-technical stakeholders
  • Comfortable being part of an on-call rotation


About Nextpoint

Nextpoint is a software company that provides cloud-based eDiscovery and litigation management solutions. The company was founded in 2001 and is headquartered in Chicago, Illinois. Nextpoint's software is used by law firms, corporations, and government agencies to manage the discovery process in legal cases. The company's solutions include document review, data processing, and trial preparation tools. Nextpoint is committed to providing its customers with secure and reliable software, and has received numerous awards for its work in this area.
Learn more about Nextpoint
Size
50 employees
Industry
Net Income
-$1 million
Founded
2001
5 Year Trend
+10%
Revenue
$5 million

Similar Jobs

More Enterprise Technology Jobs

Find similar Site Reliability Engineer jobs: