Proficient in monitoring tools such as Grafana, Kibana, or CloudWatch.
Experience in Linux system administration is essential.
Knowledge of AWS operations, particularly with EC2 instances.
Familiarity with dataflow tools like Apache NiFi or Kafka is a plus.
Understanding of DevOps tools such as Terraform and Docker is beneficial.
Responsibilities
Blend skills across Systems Administration and Operations Support across all teams in ANDS.
Monitor and maintain application systems for data reliability and accuracy.
Track system health metrics and logs to prevent downtime and ensure performance.
Identify service degradation and escalate issues as necessary for watchfloor operations.
Perform basic Linux tasks like checking services and reviewing logs.
Manage AWS Console and CLI for cloud operations and troubleshooting.
Participate in a rotating shift schedule to provide 24/7 application support.
Benefits
Comprehensive health insurance packages.
Retirement savings plan with company matching.
Opportunities for continuous professional development.
Supportive work culture emphasizing teamwork and collaboration.
Flexible work hours with rotating shifts.
Full Job Description
Responsibilities:
This is a blended position across all teams in ANDS. The successful candidate will have skills crossing from Systems Administration, and Operations Support.
Ensure data reliability and accuracy by monitoring and maintaining application systems.• ontribute to operational success through continuous performance monitoring and proactive troubleshooting, to potentially include changes to system configuration, changes to the code base, etc
Maintain system health by tracking metrics, logs, dashboards, alerts, and application status to prevent downtime and ensure optimal application performance
Support watchfloor operations by identifying service degradation, failed dataflows, system errors, and application issues, and escalating as needed.
Perform basic Linux system administration tasks, including checking service status, starting, stopping, and restarting services, reviewing logs, validating disk, memory, and CPU usage, and supporting server reboots.
Support cloud-based operations by using the AWS Console or AWS CLI to check instance health, review system status, restart or reboot servers, and assist with basic operational troubleshooting.
Work in a rotating shift schedule, 6AM-6PM / 6PM-6AM, to provide 24/7 application support on our watch floor.
Required:
Active and current TS.SCI w FSP through MD
6 YOE
Prospective candidates must have proficiency in ONE or more of the following skillsets / technologies:
Monitoring / Watchfloor Operations: Monitoring tools, such as Grafana, Kibana, Splunk, or CloudWatch.
Systems Administration: Working in a Linux environment, ability to check system health, start, stop, and restart services, review logs, and reboot servers safely.
Cloud / AWS Operations: Experience supporting EC2 instances, service health checks, and basic operational troubleshooting.
Dataflow / Application Support: Data pipeline / dataflow tools such as Apache NiFi, Cribl, Kafka, Logstash, or similar; experience monitoring application health, failed jobs, queues, data movement, and ingestion issues