Data Engineer - Top Secret Clearance

Metric5

$140K — $170K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years experience in real-time data streaming architectures and CDC patterns.
  • Proficient in Apache Kafka and skillful in Apache Flink stream processing.
  • Extensive knowledge of OpenSearch, including cluster management and search query optimization.
  • Experience with Debezium for CDC and Datadog Vector for observability.
  • Strong familiarity with Kubernetes/OpenShift for application deployment and management.
  • Ability to jointly design data contracts with software engineers, especially Node.js/NestJS.
  • Self-starter with excellent problem-solving and communication skills.

Responsibilities

  • Design and maintain Change Data Capture (CDC) pipelines with Debezium and Apache Kafka.
  • Develop and optimize Apache Flink applications for data processing.
  • Manage Vector pipelines for data processing into OpenSearch.
  • Architect and implement OpenSearch data management workflows.
  • Leverage AI to enhance data enrichment and implement vector search capabilities.
  • Deploy data infrastructure on OpenShift using GitOps practices.
  • Monitor performance across the data streaming lifecycle for reliability.

Benefits

  • Top Secret clearance required for role indicates specialized and secure project work.
  • Opportunity to work with cutting-edge technologies like AWS, OpenShift, and Apache Kafka.
  • Hands-on experience with AI/ML in data environments adds significant career growth potential.
  • Collaboration with diverse teams, enhancing cross-functional skills and knowledge.
  • Support for continuous learning and certification opportunities in advanced technologies.
Full Job Description
Location: Springfield, VA (On-site ~2 days a week)
Clearance: Top Secret Required


Position Overview
As a Data Engineer, you will work closely with O&M, development, product, design, and client teams to maintain, enhance, and deliver secure data infrastructure and search capabilities across client domains. The role requires experience building scalable, real-time data streaming architectures, designing CDC pipelines, and ensuring data is accessible, reliable, and mission-ready. You will help build, modernize, and sustain data workflows using AWS Cloud, OpenShift, and related technologies while supporting secure deployment, system performance, and continuous improvement.
Responsibilities
• Design, build, and maintain robust Change Data Capture (CDC) pipelines extracting data from MySQL databases using Debezium and routing to Apache Kafka.
• Develop and optimize Apache Flink applications to consume Debezium topics, perform complex multi-table joins, and output denormalized records back to Kafka.
• Configure and manage Vector pipelines to consume Flink-processed Kafka topics, perform object remapping, and reliably sink data into OpenSearch.
• Architect OpenSearch data management workflows, including the design and implementation of custom ingest pipelines, index templates, and lifecycle policies.
• Leverage AI skills to enhance data enrichment processes, implement vector/semantic search capabilities within OpenSearch, and support advanced analytics.
• Deploy, scale, and maintain data infrastructure on OpenShift using Helm charts and ArgoCD following GitOps best practices.
• Monitor system health, tune performance across the entire data streaming lifecycle (Kafka, Flink, OpenSearch), and ensure high availability and fault tolerance.
• Collaborate with backend engineers to ensure OpenSearch indexes are highly optimized for performant querying by downstream NestJS applications.
• Perform root cause analysis for system, application, data pipeline, and end-user issues, assisting Tier 2 support teams with complex problem resolutions.
• Maintain technical documentation related to data architecture, data flows, APIs, system configurations, and operational procedures while supporting system security coordination.
• Support occasional after-hours and weekend work for operational issue resolution, production deployments, data migrations, and maintenance windows.
Required Skills
• 7+ years of experience building scalable, real-time data streaming architectures and CDC patterns.
• Hands-on experience with Apache Kafka and deep proficiency in writing complex stream processing jobs using Apache Flink.
• Extensive experience with OpenSearch (or Elasticsearch), including cluster management, writing ingest pipelines, managing index templates, and writing complex, optimized search queries.
• Applied knowledge of integrating AI/ML models into data pipelines, working with vector databases (e.g., OpenSearch k-NN), or building AI-driven data products.
• Experience with Debezium for CDC and Vector (by Datadog) for observability and data routing.
• Proven experience deploying applications in Kubernetes/OpenShift environments with strong familiarity with infrastructure-as-code and deployment workflows using Helm and ArgoCD.
• Ability to work closely with software engineers (particularly those using Node.js/NestJS) to define data contracts and query patterns.
• Ability to design, develop, and operate highly available data services across availability zones and regions.
• Self-starter with strong problem-solving, analytical, decision-making, and verbal and written communication skills.
Preferred Skills:
• Experience working with cloud platforms such as AWS.
• Support for data quality, data validation, metadata management, and data governance practices.
• Familiarity with secure software development practices, vulnerability remediation, access control, and compliance requirements in federal or classified environments.
• Familiarity with Agile, Scrum, SAFe, or other iterative development methodologies, with experience delivering solutions to government customers.
• Experience serving in an "on-call" role supporting emergency response to application or system issues on occasion.

Certifications:
• Security+ certification is preferred.
• Other relevant certifications include CCNA, CCNP, CISA, CISSP, and CISM.

Years of Experience: 5 years+
Education: Bachelors Degree

Salary: $140,000 - $170,000

Similar Jobs

More Jobs at Metric5

More Aerospace & Defense Jobs

Find similar Data Engineer - Top Secret Clearance jobs: