We are looking for an engineer who builds the data and interface layer that makes complex software products programmatically understandable. Many powerful products expose rich but heterogeneous surfaces - source code, scripting APIs, file formats, data structures, and workflow logic. This role builds pipelines that extract and transform those surfaces into structured, queryable knowledge - designing the schemas, constructing the knowledge graphs, and exposing them through clean, typed programmatic interfaces for downstream consumption.
Job Responsibilities- Build ETL/ELT pipelines that extract data from source code, APIs, file formats, and documentation and load it into a structured knowledge store.
- Design and maintain schemas and semantic data models capturing entities, relationships, and capabilities.
- Construct and maintain knowledge graphs over heterogeneous product data.
- Develop source and metadata parsers (including source-code/AST parsing) to extract structure automatically.
- Build typed programmatic interfaces and data-access layers over the knowledge layer.
- Implement retrieval and indexing layers (e.g., embeddings, RAG) over product knowledge.
- Work with domain engineers to decompose complex product workflows into discrete, callable operations.
- Assess data sources for coverage, quality, and schema completeness across multiple products.
Job Qualifications- BS/MS in Computer Science, Mechanical Engineering, or similar.
- Strong Python; experience building and consuming REST APIs.
- Experience building data pipelines (ETL/ELT) over structured and unstructured data.
- Familiarity with graph databases and/or semantic/ontology modeling (RDF, OWL, property graphs, or equivalent).
- Experience with at least one agent framework (LangChain, LangGraph, AutoGen, CrewAI, or similar).
- Understanding of how LLMs consume context and call tools (retrieval, RAG, embeddings).
- Exposure to CAE/FEA/CFD or a related physical-simulation or engineering domain.
- Comfortable working within unfamiliar or undocumented codebases.
- Systems thinker - able to decompose a complex legacy workflow into discrete, callable steps.
Additional Skills/PreferencesNice to have:
- Vector databases.
- Data-access and API interface development.
- Parsing structured file formats.
- Surrogate modeling or related numerical methods.
Deliberately not required:
- Deep or specialist domain expertise beyond working familiarity - domain engineers provide that.
- No PhD or ML research background required.
Additional Information- Works across multiple products, building structured knowledge and interfaces over their capabilities.
- Collaborates closely with domain engineers who provide subject-matter expertise.
- Works with data pipelines, graph databases, and product API surfaces.
- Travel is not an expectation for this role. Occasional travel may occur for broad team alignment workshops, but these are infrequent.