Graph Data Engineer

Marathon TS

• $140K — $170K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's Degree with 5+ years of relevant professional experience or equivalent.
  • Active TS/SCI Clearance required.
  • Proficiency in programming and scripting for automation such as Python or Java.
  • Experience with relational database querying (e.g., SQL).
  • Hands-on experience building production data pipelines or ETL processes.
  • Familiar with at least one enterprise graph database platform, including schema design and optimization.
  • Ability to communicate complex technical concepts to both technical and non-technical stakeholders.

Responsibilities

  • Design, build, and deploy automated pipelines for metadata ingestion from various environments.
  • Lead alignment of data elements to the Enterprise Core Ontology while preserving local naming conventions.
  • Document data lineage chains and maintain provenance metadata standards.
  • Assess integration impacts on the broader Enterprise Semantic Map for effective use case alignment.
  • Maintain and optimize the Enterprise Semantic Map for efficient data access across human and AI agents.
  • Guide and mentor junior engineers on graph modeling and query construction.
  • Document schema decisions and advocate for technical positions to stakeholders.

Benefits

  • Collaborative and mission-focused work environment.
  • Opportunities for professional growth and advancement.
  • Support for continuous learning and skill development.
  • Engagement in cutting-edge technologies within the DoD sector.
Full Job Description
Graph Data Engineer
Pentagon or Reston VA
Clearance: Active TS/SCI Required
($140,000 - $170,000)


Marathon TS is seeking an experienced Graph Data Engineer to support a mission-focused data and AI initiative within a Department of Defense (DoD) environment. The ideal candidate will help design, build, and scale enterprise data capabilities that make complex information more discoverable, understandable, trusted, and usable by analysts, applications, and AI systems.

The Graph Data Engineer will develop and maintain an enterprise semantic data environment, leveraging ontology, metadata, graph databases, APIs, data integrations, and automated workflows.

This role will focus on connecting disparate data sources and creating the underlying data structures and integrations needed to support advanced search, discovery, analytics, and agentic AI workflows.

Key Responsibilities

1. Automated Source Discovery & Metadata Ingestion (Technical Metadata)
  • Supplying the "Raw Ingredients" for the Semantic Knowledge Graph: Design, build, and deploy automated pipelines that programmatically Client enterprise data assets and interface with existing data catalogs. Scan, catalog, and ingest technical metadata - including schemas, tables, columns, and API endpoints - from legacy, cloud, and distributed environments to establish baseline assets for alignment to the Enterprise Core Ontology.
    • Scaling the Semantic Map: Establish the automated pipelines and orchestrated workflows that ingest metadata at scale, replacing manual, field-by-field mapping. Own the practices that keep the ontology current as a dynamic, living "semantic control plane" rather than a static document.
    • Establishing the Entry Point for Lineage: Define how the technical origin of ingested data is registered and how metadata is captured at the point of ingestion, creating the foundation for automated provenance chains that track where data originated and how it changes over time.

2. Semantic & Provenance Mapping (Semantic & Lineage Metadata)
    • Ontological Alignment: Lead the alignment of discovered data elements from local systems to the shared Enterprise Core Ontology and specialized Domain Ontologies, with particular attention to compatibility with established institutional frameworks (e.g., DIA's DIKEM). Preserve local naming conventions while establishing standardized, shared meaning, and resolve modeling conflicts as they arise.
    • Lineage Tracking: Design and maintain data lineage chains within the Provenance Layer, applying industry lineage standards to document where data originates, how it is transformed, and who governs it.
    • Graph Querying & Validation: Write, optimize, and review graph queries supporting metadata retrieval, logical validation, and graph manipulation. Establish reusable query patterns and validation checks the wider team can build on.

3. Enterprise Systems Thinking & Alignment
    • Big-Picture Integration: Assess how newly integrated data sources and automated pipelines affect the broader Enterprise Semantic Map, selected use cases, downstream consumers, and enterprise search and discovery - and adjust the design accordingly.
    • Downstream Enablement: Connect data assets to relevant mission metadata so technical capabilities can be clearly linked to the mission workflows they support.
    • Governance Compliance: Ensure enterprise assets are associated with appropriate governance metadata, including ownership, classifications, handling rules, and access constraints. Translate complex data policies into machine-readable semantic structures.

4. Smart Search & Agent Enablement
    • Semantic Control Plane Ownership: Maintain and optimize the Enterprise Semantic Map within enterprise graph database platforms so human analysts, applications, and autonomous AI agents can efficiently search, navigate, and Client resources. Tune schema and query performance as the graph grows.
    • Agent Integration: Partner with AI engineers so planning, research, and tool agents can dynamically query the graph, and help define the grounded, trustworthy reasoning and retrieval strategies those agents depend on.

5. Technical Leadership & Mentorship
    • Mentorship: Guide junior engineers on graph modeling, query construction, and pipeline development, and review their work.
    • Design Documentation & Advocacy: Document schema decisions, modeling rationale, and runbooks so the design is reproducible, and represent technical positions clearly to architects, program leadership, and government stakeholders.

Required Experience/Clearance
  • Bachelor's Degree with 5+ of relevant professional experience or equivalent.
  • Active TS SCI Clearance.
  • Core Technical Skills: Foundational proficiency across the following areas, demonstrated in any comparable technology:
    • Programming and scripting for automation (e.g., Python, Java, or a comparable general-purpose language)
    • Relational database querying (e.g., SQL)
    • Structured and semi-structured data formats (e.g., JSON, XML, YAML)
    • Graph query languages for retrieval, validation, and manipulation (e.g., Cypher for property graphs, SPARQL for RDF/triple stores)
    • Knowledge graph concepts, including nodes, edges, relationships, and metadata schemas
  • Data Engineering Experience: Hands-on experience building and operating production data pipelines or ETL (Extract, Transform, Load) processes, including error handling, monitoring, and scheduling.
  • Graph Database Platforms: Practical experience with at least one enterprise graph database platform, including schema design and query performance considerations.
  • API & Systems Integration: Experience integrating heterogeneous systems through APIs across legacy, cloud, and distributed environments.
  • Systems-Thinking Mindset: Ability to reason about how individual pipelines and modeling choices propagate through a broader enterprise ecosystem, and to weigh trade-offs explicitly.
  • Attention to Detail: Precision in aligning metadata terms, formatting data endpoints, and maintaining technical schemas.
  • Communication & Stakeholder Engagement: Ability to explain semantic and architectural decisions to both engineering peers and non-technical mission stakeholders, and to document them durably.

Preferred Qualifications
  • Ontology & Semantic Standards: Working experience with formal ontology or semantic web standards (e.g., RDF, OWL, SHACL) and with established government- or defense-related semantic models.
  • Agentic AI & AI Frameworks: Experience with LLM orchestration, retrieval-augmented generation, or agentic workflows, particularly where a graph provides grounding.
  • Data Lineage & Metadata Standards: Applied experience with open lineage specifications or metadata management frameworks.
  • Data Catalogs & Stewardship: Experience with metadata catalog environments and data stewardship systems.
  • Workflow Orchestration: Experience with pipeline scheduling and orchestration tooling.
  • Cloud & Deployment: Familiarity with cloud data platforms, containerized deployment, and CI/CD practices.
  • Mission Domain Exposure: Prior experience supporting defense, intelligence community, or other regulated enterprise data environments.


#CJJOBS

Similar Jobs

More Jobs at Marathon TS

More Aerospace & Defense Jobs

Find similar Graph Data Engineer jobs: