AI Data Platform Engineer

Banner Quality Management Inc.

• $110K — $130K *
Aerospace & Defense
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's in Computer Science, Data Science, AI, Engineering, or related field (PhD preferred).
  • 3+ years experience in building and deploying AI/data-intensive systems in production.
  • Proficiency in Python and SQL, with knowledge of data engineering libraries.
  • Hands-on AWS experience across multiple services (storage, compute, AI).
  • Experience designing data lake architectures and data processing pipelines.
  • Knowledge of API design, authentication, and containerization (Docker).
  • Understanding of data handling in regulated domains (CUI/PII-aware design).

Responsibilities

  • Architect and operate a cloud data lake on AWS for various data types.
  • Design end-to-end data pipelines with orchestration and monitoring.
  • Develop retrieval-augmented generation (RAG) and agentic GenAI capabilities.
  • Document and secure APIs for accessing datasets and enterprise systems.
  • Implement authentication for AI services using OIDC/OAuth 2.0.
  • Build evaluation frameworks for GenAI systems and monitor performance.
  • Integrate AI capabilities into existing applications and collaborate with experts.

Benefits

  • Work within a new and dynamic team focused on AI innovations.
  • Opportunity to have a direct impact on customer safety and mission assurance.
  • Engage with cutting-edge technologies in a high-stakes environment.
  • No relocation support implies proximity to a dedicated local talent pool.
  • Exposure to federal guidelines and governance frameworks in AI.
Full Job Description
AI Software Developer/Engineer

BQMI has a rewarding position as an AI Data Platform Engineer. This is a new position and new capabilities within an established development team. The Customer is seeking an experienced team member who will design, build, deploy, and operationalize GenAI and data platform capabilities that implement and improve Customer Safety & Mission Assurance (SMA) use cases for the Organizational Safety & Mission Assurance Division. The Customer's primary infrastructure is AWS. The team member will architect and grow a data lake that unifies structured, semi-structured, and unstructured sources (relational databases, documents, meeting transcripts, news/RSS feeds, M365/SharePoint content), expose those sources to AI workflows through secure APIs and Model Context Protocol (MCP) servers, and integrate with existing applications and enterprise identity, while meeting Customer policy, security, accessibility, and governance requirements.

Essential Duties & Responsibilities:

  • Architect, build, and operate a cloud data lake on AWS that ingests and organizes structured, semi-structured, and unstructured data from relational databases (e.g., SQL Server, managed RDS), object storage, documents, transcripts, RSS/news feeds, and M365/SharePoint sources, and makes that data AI-ready.
  • Design and build end-to-end data pipelines: ingestion, transformation, document parsing and chunking, metadata enrichment, embedding generation, and loading into vector/semantic and structured stores, with orchestration, scheduling, data quality checks, and monitoring.
  • Develop and operationalize retrieval-augmented generation (RAG) and agentic GenAI capabilities using AWS-native foundation model services and open frameworks, including knowledge bases, embeddings, vector search, natural-language-to-SQL, and tool use over the various datasets.
  • Design, develop, secure, and document APIs and MCP servers that expose lake datasets and enterprise systems to LLM orchestrators, chat interfaces, and web applications as governed, discoverable tools.
  • Implement authentication and authorization for AI services, APIs, and data access using OIDC/OAuth 2.0 and Microsoft Entra ID; configure and administer identity providers (including Keycloak) for scopes, roles, service-to-service authentication, and federation.
  • Build evaluation and observability for GenAI systems: retrieval quality, answer groundedness, hallucination detection, validation of NL-to-SQL outputs against source data, tracing, and cost/latency monitoring.
  • Embed data governance into pipelines and retrieval: data classification and marking, document-level and row/column-level access controls, PII handling, lineage, and audit logging.
  • Design and optimize prompts and agent workflows for large language models (LLMs) to ensure accurate, grounded, context-aware outputs with source citations.
  • Evaluate and recommend tools: help the team select the right AWS services, open-source, and third-party components for each job based on capability, cost, security accreditation, and maintainability, avoiding over-engineering.
  • Integrate AI capabilities into existing web applications, tools, training platforms, and M365/SharePoint solutions.
  • Collaborate with subject-matter experts, engineers, data stewards, and mission programs to define use cases, gather requirements, and measure impact (e.g., reduced analysis time, improved risk foresight, higher detection accuracy).
  • Ensure AI systems follow the Customer's "AI advises, humans decide" posture and responsible AI principles: human-in-the-loop review, source traceability, clear distinction between retrieved data and AI-generated content, explainability, robustness, safety, and compliance with federal guidelines (e.g., NIST AI Risk Management Framework, Customer AI governance).
  • Stay current on frontier AI advances (e.g., agentic systems, the MCP ecosystem, trustworthy AI) and adapt them to constrained, high-assurance aerospace domains.
  • Document architecture, pipelines, code, and processes to support reproducibility, audits, and knowledge transfer across the Customer organization.
  • Must live within Greater Cleveland Area.
  • No relocation funding available


Essential Skills:

  • Bachelor's or Master's degree in Computer Science, Data Science, Artificial Intelligence, Engineering, or a related field (PhD preferred for senior levels).
  • 3+ years of hands-on experience building and deploying AI and data-intensive systems in production.
  • Proficiency in Python and strong SQL; experience with modern data engineering libraries and frameworks.
  • Hands-on AWS experience across storage, compute, data, and AI services (object storage, managed relational and vector databases, serverless and containerized compute, managed foundation model and knowledge base services), including IAM and security best practices.
  • Demonstrated experience designing data lake or lakehouse architectures and pipelines for heterogeneous data at scale, including document processing and embedding workflows.
  • Experience building and securing REST APIs and/or MCP servers: API design, authentication, versioning, containerization (Docker), and deployment.
  • Working knowledge of OIDC/OAuth 2.0 flows, JWTs, and enterprise identity providers (Microsoft Entra ID required; Keycloak a plus).
  • Strong understanding of data handling in regulated/high-stakes domains: CUI/PII-aware design, access controls, auditability, and NLP over technical documents.


Experience/Education:

  • Experience implementing GenAI solutions (not just experimentation) using APIs, SDKs, and frameworks (e.g., Amazon Bedrock, LangChain/LangGraph, LlamaIndex, LiteLLM or other model gateways, Hugging Face).
  • Work with large language models (RAG, embeddings and vector databases, agent frameworks, NL-to-SQL) for technical document analysis or knowledge extraction.
  • Experience with data workflow orchestration and processing tools (e.g., AWS Step Functions, Glue, Lambda, Airflow, dbt, Spark); specific tooling is not prescribed and candidates are expected to help select it.
  • Experience evaluating and monitoring LLM applications (e.g., RAG evaluation frameworks, LLM-as-judge, tracing/observability platforms, OpenTelemetry).
  • Experience implementing robust guardrails, content safety, and groundedness controls.
  • ML literacy sufficient to operationalize, monitor, and retrain existing models and embedding pipelines; familiarity with scikit-learn or PyTorch a plus.
  • C#/.NET a plus, for integration with existing web applications and backend services.
  • TypeScript/Node.js a plus, for MCP servers and frontend integration.
  • Experience integrating with M365/SharePoint (e.g., Microsoft Graph API).
  • Experience with Git/GitHub workflows, branching strategies, and code review in an Agile team.
  • Prior Federal government, defense, or contracting experience


Personality or self-management skills:

  • Strong communication, presentation, and storytelling skills
  • Experience collaborating with multiple cross-functional partners
  • Experience working in a fast-paced, iterative environment where multi-tasking and time-management skills are critical
  • Thrive under pressure and enjoy working in an environment with competing priorities
  • Experience supporting developers in the implementation of design deliverables
  • Ability to perform work within specific timeframes and adhere to deadlines
  • Willingness to keep skills current and quickly adapt to new technologies, as needed.


Similar Jobs

More Jobs at Banner Quality Management Inc.

  • AI Data Platform Engineer
    $110K — $130K *
    Cleveland, OH 44130 (Cuyahoga County)
    Aerospace & Defense
    In-Person
  • AWS Administrator
    $95K — $115K *
    Cleveland, OH 44130 (Cuyahoga County)
    Information Technology
    In-Person

More Aerospace & Defense Jobs

Find similar AI Data Platform Engineer jobs: