Comfortable with the Linux environment or UNIX CLI
Experience with programming or scripting languages
Experienced in creating processes, procedures and SOP documentation
General understanding of TCP/IP, HTTP, and related protocols
Responsibilities
Lead the platform operations team in transitioning legacy systems to a modern DevOps platform
Identify and resolve operational problems in a micro-service environment
Collaborate with developers to address deployment and runtime challenges
Perform analysis and debugging across multiple technologies
Prioritize issues to maintain applications within error budgets and SLOs
Document processes and Standard Operating Procedures (SOPs) with team input
Compile postmortems and action items for minimizing future outages
Conduct interviews and assist in hiring new team members
Train new members and provide support for technical issues
Offer on-call support to internal developers and staff
Benefits
Flexible working hours
Remote work options
On- and off-site training courses
Conference attendance
Tuition reimbursement
Full Job Description
Overview
This opportunity is full time at the NCBI in Bethesda, MD and/or remote.
The Systems & DevOps team is responsible for the efficient operation of infrastructure to run NCBI's many applications. This includes providing convenient, scalable solutions for development, deployment, and operations across teams, languages, and cloud and on-prem environments.
This is a great opportunity to work on challenging problems in a technical, scientific, and goal-oriented environment. NCBI offers flexible working hours, remote options, on- and off-site training courses, and conference attendance and tuition reimbursement.
Duties & Responsibilities
NCBI has built a modern DevOps platform based on GitLab and Kubernetes, and is looking to create a team of support engineers to assist internal developers with transitioning legacy software development and deployment to the new DevOps platform. You would be the leader of our platform operations team.
Identify and resolve operational problems in a micro-service environment
Work with developers to resolve deployment and runtime problems
Perform analysis and debugging work across multiple technologies
Prioritize issues to keep applications within error budgets and meeting their SLOs
Provide technical solutions to a wide range of problems and user requests
Document process, procedures and SOPs by soliciting feedback and suggestions from team members
Compile postmortems and action items to minimize future outages
Interview other people for team member roles, and decide which ones to recommend for hire.
Train new team members, and assist them with issues.
Provide on-call support to NCBI's internal developers and other staff.
Requirements
BS degree in STEM or equivalent experience
Customer-focused, team-oriented disposition
Good systems debugging skills
Comfortable with the Linux environment or UNIX CLI
Experience with some programming or scripting language
Have experience creating processes, procedures and SOP documentation
General understanding of TCP/IP, HTTP, and related protocols
Initiative to take ownership of tasks and drive them to completion
Comfortable dealing with users with varying levels of IT knowledge
Eager to learn new technologies
Strong communication and soft skills to interface with customers, peers and management
Good judgement, sense of integrity, and responsibility
Preferred Experience / Skillsets
Kubernetes, OpenShift, Cloud or Linux experience
Experience with:
Service Reliability Engineering in any capacity
Linux systems administration
Automated CI servers, especially TeamCity and/or GitLab
Automation programming/scripting in any of: bash, Ruby, Python, Go, Java, Scala, Rust, C++, Perl
Automated configuration management, such as Puppet, Ansible, Chef, bcfg2, cfengine, etc. Puppet is preferred.
Version control systems, especially git
Service Mesh technologies (e.g., linkerd, Istio)
Configuring or using monitoring and alerting technologies (TIGK stack, Grafana, Prometheus, OpsGenie)
Confluence, Jira, and Microsoft Office suite
GitOps tools, especially ArgoCD
Google Anthos
Understanding of:
Linux internals (system calls, file systems, processes, etc.)
Linux network configuration
Linux application containerization, especially Docker
Attached network storage technologies
Cloud computing environment such as AWS, GCP or Azure