Full Job Description
You are an experienced Cloud Ops Engineer who thrives in a fast-paced and AI-forward environment. You have a passion for innovation, solid design principles, and high-quality development. You excel at designing, improving, and maintaining secure, performant infrastructure and enjoy automating processes to ensure efficiency and scalability. You are a proactive problem solver with a strong understanding of system and networking concepts.
What You'll Do:
• Infrastructure Design and Maintenance:
- Design, improve, and maintain secure, durable, and performant infrastructure to power APIs, AI operations, web applications, and data mining/ETL workflows to meet established SLAs.
- Collaborate with developers to bring new products and services into production.
• Automation and Monitoring:
- Automate testing, deployment, and monitoring of all products and services throughout the software development lifecycle.
- Continuously improve operational processes and apply best practices to ensure scalability, security, and availability.
• Security and Compliance:
- Proactively meet standards for information security and compliance, such as SOC 2/ISO27001/CMMC.
- Implement and uphold security measures across all infrastructure components.
Requirements:
• Professional Experience:
- At least 5 years of professional experience in a Cloud Ops / Platform Ops / DevOps role maintaining production infrastructure, preferably supporting a highly available environment for a SaaS or cloud service provider.
• Technical Proficiency:
- Strong working knowledge of AWS services such as EC2, ECS or EKS, Lambda, API Gateway, RDS, DynamoDB, Cloudwatch, S3, Code/Build/Pipeline/Deploy, VPC Lattic etc.
- Strong working knowledge of Terraform or similar tools, Ansible, AWS CLI/SDK, Boto.
- Proficiency with scripting languages such as Python, Bash, etc., and Linux environments.
- Strong understanding of system and networking concepts and troubleshooting techniques for bare metal and containerized workloads.
- Experience supporting AI and ML systems, agent orchestration frameworks (e.g., AgentCore or similar), and experience integrating LLMs into production systems.
• Additional Skills:
- Experience with release automation, system administration and configuration, and system debugging.
Nice to Have
• Databricks or Snowflake infrastructure experience.
• Cost optimization for AI and LLM workloads.
• Internal developer platform experience.
Base Salary Range: $146,000 - $190,000
The salary range reflects the expected base compensation for a fully qualified candidate at this level based on experience, qualifications, and market data at the time of posting.
U.S.-Based Benefits + Perks (for Full Time Employees):
At SpyCloud, we are committed to working alongside individuals who are equally passionate about preventing cybercrime, regardless of their department or role. Guided by our core values in all business decisions, we prioritize unity in our mission and ensure all SpyCloud employees have the support and benefits they need to stay focused on our goals. In addition to our engaging workspace in South Austin, flexible and remote-friendly work options, and competitive salary package, we offer our employees a comprehensive benefits package that includes:
• 401(k) with Employer Contribution
• Health, Vision, and Dental Insurance
- Health Savings Account (HSA) available with Employer Contribution
• Employer Paid Life, Short-term, and Long-term Disability Insurance
• Generous PTO Plan and 16 paid holidays per year
U.K.-Based Benefits + Perks (for Full Time Employees):
• Retirement Savings Plan with Employer Contribution
• Employer Provided Private Health Insurance and Healthcare Cashplan
• Employer Paid Life Insurance, Income Replacement, and Critical Illness Protection
• Employee Assistance Program
• 25 days annual accrued holiday plus 8 paid bank holidays and a paid company shutdown during winter holidays