Full Job Description
Reliability Engineer
Responsible for the availability, performance, monitoring, and incident response, among
other things, of the cloud platforms and services. Ensure that everything that goes to
production complies with a set of general requirements like diagrams, dependencies of
other services, monitoring and logging plans, backups and possible high availability setups.
Manages uncaught exceptions, hardware degradation, networking problems, high usage of
resources, or slow responses that could happen at any time. Uses metrics such as mean
time to recover (MTTR) and mean time to failure (MTTF). Considered an emerging authority,
who applies extensive technical expertise. Develops technical solutions to complex
problems. Exercises considerable latitude in determining objectives and approaches to
Assignment.
Certifications: IAT-2, 1 or more cloud certifications
Minimum Education Requirement: Bachelor's degree
Minimum/General Experience: 8 years of experience
Location: Remote (preferred located in Boston)
Clearance: Secret