Cambridge Mobile Telematics

Principal Site Reliability Engineer, Machine Learning

Cambridge Mobile Telematics$142K — $177K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree or equivalent experience in a related field
  • 7+ years in Site Reliability Engineering or IT
  • Ability to design and document systems and automate processes
  • Intermediate to expert experience with AWS services (EC2, EKS, etc.)
  • Intermediate to expert monitoring experience with CloudWatch and Datadog
  • Intermediate to expert skills in maintaining AWS compute services for ML workloads
  • Proficiency in at least one programming language, primarily Python
  • Intermediate to expert experience with Terraform and other Infrastructure as Code tools
  • In-depth knowledge of Linux OS and experience with Docker/Kubernetes
  • Experience leading system design projects and architecture changes

Responsibilities

  • Own SLOs and operational health of Ray clusters on AWS EKS and Databricks
  • Ensure uptime and scalability of workloads using CloudWatch and Datadog
  • Tune EKS Ray workloads for efficiency and resilience
  • Manage AWS Databricks including administration and IAM roles
  • Maintain cost efficiency and visibility for EC2 and EKS workloads
  • Perform upkeep of underlying AWS infrastructure and security updates
  • Implement infrastructure-as-code with Terraform and automate tasks
  • Lead incident response and facilitate postmortem analyses
  • Address additional tasks as needed

Benefits

  • Competitive salary and annual performance bonus
  • Equity opportunities through RSUs
  • Comprehensive medical, dental, and vision insurance
  • 401k matching and parental leave options
  • Unlimited Paid Time Off with flexible scheduling
  • Work-from-home policy based on role needs
  • Summer Fridays for enhanced work-life balance
  • Participation in diverse employee resource groups
  • Access to wellness, education, and assistance programs
Full Job Description
Responsibilities:
  • Use independent judgment and discretion to own SLOs, error budgets, and the operational health of Ray clusters running on AWS EKS and Databricks workloads on AWS EC2 across multiple accounts and regions
  • Maintain the observability of uptime, availability, and scalability of EKS Ray and Databricks workloads using CloudWatch and Datadog, including defining alerting that maps to SLOs
  • Operate and tune EKS Ray workloads at scale including autoscaling, GPU scheduling, and automated failure recovery
  • Manage Databricks on AWS including workspace administration, cluster policies, Unity Catalog, job orchestration, and IAM Roles and Policies
  • Maintain ongoing cost visibility, cost optimization, and capacity planning across EC2 and EKS workloads, including through the use of On Demand Capacity Reservations and Spot lifecycle
  • Perform ongoing maintenance of the underlying EC2 and EKS infrastructure, including regular security updates and operating system upgrades
  • Codify everything as infrastructure-as-code using Terraform and CI/CD pipelines, enabling updates through Pull Requests with approval workflows, while also automating maintenance tasks to reduce toil
  • Lead incident response for Data Science and Machine Learning platform outages, run blameless postmortems, and drive systemic remediation, including participating in an on-call rotation
  • Complete any additional tasks as they arise

Qualifications:
  • Bachelor's degree or equivalent years of experience and/or certification in a related field
  • 7+ years working in Site Reliability Engineering or Information Technology
  • Design and document systems, including writing and reviewing code, to automate away problems within your team's domain
  • Intermediate to expert experience deploying and maintaining AWS services such as EC2, ECS, EKS, SQS, Lambda, Dynamo, RDS/Aurora, S3, and IAM
  • Intermediate to expert experience monitoring services and applications using tools such as CloudWatch Metrics, CloudWatch Logs, and Datadog, including defining and configuring alerts and SLO reports
  • Intermediate to expert experience maintaining the uptime and scalability of AWS compute services used for Machine Learning and Data Science workloads, specifically EC2 and EKS
  • Intermediate to expert coding skills in at least one programming language; we work primarily in Python
  • Intermediate to expert experience using Infrastructure as Code platforms and CI/CD pipelines, specifically Terraform, to manage AWS infrastructure and services
  • In-depth knowledge & experience with Linux operating systems (Amazon Linux, Ubuntu) on EC2 and Docker / Kubernetes
  • Experience with leading projects in system design, architecture changes, and technology selection

Compensation and Benefits:
  • Fair and competitive salary based on skills and experience, and annual performance bonus
  • Equity may be awarded in the form of Restricted Stock Units (RSUs)
  • Medical, Dental, Vision and Life Insurance, matching 401k, short-term & long-term disability and parental leave
  • Unlimited Paid Time Off including vacation, sick days & public holidays
  • Flexible scheduling and work from home policy depending on role and responsibilities

Base Salary Range
  • The base salary range for this position is: $142,000 to $177,600. This range is specifically for Cambridge, MA

Additional Perks:
  • Work on a mission with real impact: crashes prevented, injuries avoided, lives protected around the world
  • Join an industry leader - 65 million drivers protected, powering 140+ programs across 25 countries
  • Be part of the team inventing the future of mobility and road safety
  • Move fast, own outcomes, do work that matters
  • High ownership, small teams, and direct access to leadership - no layers between your work and its impact
  • Unlimited PTO, flexible scheduling, competitive salary, annual performance bonus, RSUs, and full benefits including medical, dental, vision, and 401k match
  • Summer Fridays provide team members with half days to recharge
  • Join one of our employee resource groups: Black, AAPI, LGBTQIA+, Women, Book Club, and Health & Wellness
  • Comprehensive wellness, education, and employee assistance programs

About Cambridge Mobile Telematics

Cambridge Mobile Telematics (CMT) is a technology company that provides mobile telematics and analytics solutions for insurers, rideshares, and fleets. The company's platform uses sensors and mobile applications to collect data on driving behavior, which is then analyzed to provide insights into risk and safety. CMT's solutions are used by insurance companies to offer usage-based insurance (UBI) policies, by rideshare companies to monitor driver behavior and improve safety, and by fleets to optimize operations and reduce risk. The company was founded in 2010 and is headquartered in Boston, MA.
Learn more about Cambridge Mobile Telematics
Size
201 employees
Industry
Founded
2010

Similar Jobs

More Jobs at Cambridge Mobile Telematics

More Information Technology Jobs

Find similar Principal Site Reliability Engineer, Machine Learning jobs: