Data Engineer - Top Secret Clearance

Metric5

$140K — $170K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of experience in real-time data streaming architectures and CDC patterns
  • Hands-on experience with Apache Kafka and expert in Apache Flink stream processing
  • Extensive knowledge of OpenSearch, including cluster management and optimized query writing
  • Experience integrating AI/ML models in data pipelines and working with vector databases
  • Familiarity with Debezium for CDC and observability through Vector
  • Proven deployment experience in Kubernetes/OpenShift with Helm and ArgoCD
  • Strong collaborative skills with software engineers, particularly in Node.js/NestJS
  • Self-starter with problem-solving and communication skills

Responsibilities

  • Design and maintain Change Data Capture pipelines from MySQL to Kafka
  • Optimize Apache Flink applications for complex data processing
  • Manage Vector pipelines to incorporate Flink-processed Kafka topics into OpenSearch
  • Develop data management workflows in OpenSearch with custom pipelines
  • Enhance data enrichment processes using AI and implement vector search capabilities
  • Deploy and scale data infrastructure on OpenShift via GitOps practices
  • Monitor system performance and ensure high availability across the data streaming lifecycle
  • Assist in troubleshooting and root cause analysis for data pipeline issues
  • Document data architecture and operational procedures ensuring system security
  • Support after-hours work for operational resolutions and production deployments

Benefits

  • 100% individual Health & Dental Insurance paid by the company
  • Vision Insurance
  • Life & Short Term Disability Insurance
  • 401K with company match and immediate vesting
  • Generous paid vacation and 9 paid holidays plus 2 floating holidays
  • Parental Leave
  • Employee Bonuses
  • Professional Development Reimbursement Program
  • Tuition Assistance Program
Full Job Description
Location: St. Louis, MO (On-site ~2 days a week)
Clearance: Top Secret Required


Position Overview
As a Data Engineer, you will work closely with O&M, development, product, design, and client teams to maintain, enhance, and deliver secure data infrastructure and search capabilities across client domains. The role requires experience building scalable, real-time data streaming architectures, designing CDC pipelines, and ensuring data is accessible, reliable, and mission-ready. You will help build, modernize, and sustain data workflows using AWS Cloud, OpenShift, and related technologies while supporting secure deployment, system performance, and continuous improvement.
Responsibilities
• Design, build, and maintain robust Change Data Capture (CDC) pipelines extracting data from MySQL databases using Debezium and routing to Apache Kafka.
• Develop and optimize Apache Flink applications to consume Debezium topics, perform complex multi-table joins, and output denormalized records back to Kafka.
• Configure and manage Vector pipelines to consume Flink-processed Kafka topics, perform object remapping, and reliably sink data into OpenSearch.
• Architect OpenSearch data management workflows, including the design and implementation of custom ingest pipelines, index templates, and lifecycle policies.
• Leverage AI skills to enhance data enrichment processes, implement vector/semantic search capabilities within OpenSearch, and support advanced analytics.
• Deploy, scale, and maintain data infrastructure on OpenShift using Helm charts and ArgoCD following GitOps best practices.
• Monitor system health, tune performance across the entire data streaming lifecycle (Kafka, Flink, OpenSearch), and ensure high availability and fault tolerance.
• Collaborate with backend engineers to ensure OpenSearch indexes are highly optimized for performant querying by downstream NestJS applications.
• Perform root cause analysis for system, application, data pipeline, and end-user issues, assisting Tier 2 support teams with complex problem resolutions.
• Maintain technical documentation related to data architecture, data flows, APIs, system configurations, and operational procedures while supporting system security coordination.
• Support occasional after-hours and weekend work for operational issue resolution, production deployments, data migrations, and maintenance windows.
Required Skills
• 7+ years of experience building scalable, real-time data streaming architectures and CDC patterns.
• Hands-on experience with Apache Kafka and deep proficiency in writing complex stream processing jobs using Apache Flink.
• Extensive experience with OpenSearch (or Elasticsearch), including cluster management, writing ingest pipelines, managing index templates, and writing complex, optimized search queries.
• Applied knowledge of integrating AI/ML models into data pipelines, working with vector databases (e.g., OpenSearch k-NN), or building AI-driven data products.
• Experience with Debezium for CDC and Vector (by Datadog) for observability and data routing.
• Proven experience deploying applications in Kubernetes/OpenShift environments with strong familiarity with infrastructure-as-code and deployment workflows using Helm and ArgoCD.
• Ability to work closely with software engineers (particularly those using Node.js/NestJS) to define data contracts and query patterns.
• Ability to design, develop, and operate highly available data services across availability zones and regions.
• Self-starter with strong problem-solving, analytical, decision-making, and verbal and written communication skills.
Preferred Skills
• Experience working with cloud platforms such as AWS.
• Support for data quality, data validation, metadata management, and data governance practices.
• Familiarity with secure software development practices, vulnerability remediation, access control, and compliance requirements in federal or classified environments.
• Familiarity with Agile, Scrum, SAFe, or other iterative development methodologies, with experience delivering solutions to government customers.
• Experience serving in an "on-call" role supporting emergency response to application or system issues on occasion.

Certifications:
• Security+ certification is preferred.
• Other relevant certifications include CCNA, CCNP, CISA, CISSP, and CISM.

Years of Experience: 5 years+
Education: Bachelors Degree

Salary: $140,000 - $170,000

Our benefits include:
  • Health & Dental Insurance with 100% of individual coverage paid for by the company
  • Vision Insurance
  • Life & Short Term Disability Insurance
  • 401K with company match (employees are immediately vested)
  • Paid Vacation
  • 9 Paid Holidays per year (plus 2 paid floating holidays)
  • Parental Leave
  • Employee Bonuses
  • Professional Development Reimbursement Program
  • Tuition Assistance Program

Similar Jobs

More Jobs at Metric5

More Information Technology Jobs

Find similar Data Engineer - Top Secret Clearance jobs: