Maximus

Site Reliability Engineer (Systems Manager)

Maximus$130K — $150K *
Aerospace & Defense
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • Active Secret clearance is required.
  • U.S. citizenship is mandatory due to contract requirements.
  • Over 15 years of experience in IT platforms and infrastructure is essential.
  • Familiarity with Infrastructure as Code (IaC) tools, such as Terraform, Chef, and Ansible, is required.
  • Proficient with Jenkins and GitLab.

Responsibilities

  • Collaborate with teams to design and maintain highly available, fault-tolerant systems.
  • Monitor system performance using various tools, resolving incidents proactively.
  • Conduct capacity planning and implement optimization techniques for systems.
  • Implement and oversee backup and disaster recovery solutions.
  • Ensure application reliability and scalability in production environments.
  • Stay informed on best practices in site reliability engineering and motivate adherence to IT standards.
  • Maintain detailed incident records and produce metric reports for stakeholders.

Benefits

  • On-site position in Colorado Springs, CO.
  • A commitment to 24/7 operational support, including on-call responsibilities.
  • Opportunity for professional growth through training program development and implementation.
Full Job Description
Description & Requirements

Maximus is seeking a Systems Manager to provide expertise to a federal client in support of their mission critical systems in defense of our Homeland.

The Systems Manager will ensure the availability, reliability, and performance of critical systems and applications, utilizing a diverse set of relevant technologies, for a federal client's operations. As a crucial member of our team, the Systems Manager will play a pivotal role in designing, implementing, and managing scalable and highly available systems.

This is an on-site position based in Colorado Springs, CO requiring an active Secret clearance.

Maximus TCS (Technology and Consulting Services) Internal Job Profile Code: TCS137, T5, Band 8

Job-Specific Essential Duties and Responsibilities:
• Collaborate with cross-functional teams to design, implement, and maintain highly available, fault-tolerant systems, leveraging a range of technologies.
• Monitor and analyze system performance, utilizing a diverse set of monitoring and alerting tools, and take proactive actions to prevent and resolve incidents.
• Conduct system capacity planning and scaling, utilizing performance testing and optimization techniques across various technologies.
• Implement and maintain robust backup and disaster recovery solutions, leveraging relevant technologies, to ensure data integrity and system availability.
• Collaborate with development teams to ensure application reliability, performance, and scalability in production environments.
• Stay updated with industry best practices and emerging trends in site reliability engineering and incorporate relevant technologies. Motivate the team to adhere to IT best practices and deliver outstanding customer service and satisfaction.
• Assist in creating and maintaining a training program to increase business, customer service, and technical knowledge.
• Participate in the organization's change management process.
• Maintain detailed records of incidents, including reporting, classification, and resolution steps for daily and weekly metric reporting.
• Gather and present PMR metric reporting to management, contract leadership, and government stakeholders.
• Use monitoring tools to proactively identify and address potential issues and generate reports on incident trends.
• Collaborate with the Operations Support Center Lead to ensure continuity of daily operations.
• Other duties as assigned.

Job-Specific Minimum Requirements:
• Active Secret clearance is required.
• Due to contract requirements, only U.S. citizens can be considered.
• 15+ years of relevant experience supporting IT platforms, environments, and infrastructure is required.
• Experience in some or all the following: Infrastructure as Code (IaC) tools (Terraform, Chef, Ansible), Jenkins, GitLab
• Candidates must reside within a commutable distance for daily onsite work and on-call requirements.
• This contract supports systems that require 24x7x365 uptime. Candidates must be willing and able to meet recall requirements, including participation in a rotational on-call schedule.

#techjobs #clearance

Minimum Requirements

TCS137, T5, Band 8

Minimum Salary

$

130,000.00

Maximum Salary

$

150,000.00

About Maximus

MAXIMUS, Inc. is an American, outsourcing company that provides business process services to government health and human services agencies in the United States, Australia, Canada, Saudi Arabia, Singapore, and the United Kingdom. MAXIMUS focuses on administering government-sponsored programs, such as Medicaid, the Children's Health Insurance Program (CHIP), health care reform, welfare-to-work, Medicare, child support enforcement, and other government programs. The company is based in Reston, Virginia, has 13,000 employees and a reported annual revenue of $3.8 billion in fiscal year 2020.
Learn more about Maximus
Size
35,800 employees
Market Cap
$4.4 billion
Industry
Net Income
$219.8 million
Founded
1975
5 Year Trend
+13.6%
Revenue
$3.5 billion
NASDAQ

Similar Jobs

More Jobs at Maximus

More Aerospace & Defense Jobs

Find similar Site Reliability Engineer (Systems Manager) jobs: