Job DescriptionSummary:As a Staff Data Engineer, you will lead the design and evolution of the data platform that powers BD Connected Care's AI initiatives. You will architect and build scalable data pipelines, models, and governance solutions that deliver trusted, high-quality data to Machine Learning engineers and other consumers across the organization.
This is a hands-on technical leadership role that combines software development, data architecture, and platform strategy. You will establish engineering standards, drive data quality and governance, and make key architectural decisions that shape the long-term direction of the AI data platform. Working cross-functionally with engineering, product, and clinical teams, you will deliver scalable data solutions that advance BD's mission of improving patient outcomes through innovation and data-driven healthcare.
The ideal candidate can communicate complex data design concepts in a clear, concise manner and confidently explain trade-offs, risks, and limitations to both technical and non-technical audiences, including executive, clinical, and engineering stakeholders.
Key Responsibilities:
- Architecture and Technical Direction: Own the data architecture for the team's AI workflows, including the warehouse, lake, and serving layers. Set the technical direction and standards that engineers across the team build against, make long-term AI data platform decisions , and design for source diversity so a new device, EMR, or supply chain source can be onboarded without redefining the model.
- Ingestion and Integration:Propose, design, and implement ingestion from BD product, operational, and clinical-adjacent sources into the analytics and ML data layers, spanning device transactions, inventory, ordering, purchasing, and supply signals, in batch, streaming, and change-data-capture modes.
- Data Modeling:Design relational and analytical schemas in the cloud data warehouse and the S3-based lake for forecasting, analytics, and feature stores. Author and tune complex SQL, build dimensional and event-based models, manage incremental loads, slowly changing dimensions, and late-arriving data, and deliver analytics-ready models for downstream consumers.
- Reconciled Histories:Build continuous, reconciled records that align events across systems with different refresh cadences and timestamps, covering procurement, storage, distribution, dispense, administration, waste, and return. Capture the observation time of each value alongside the value itself.
- Data Quality and Contracts:Own data quality as an outcome. Negotiate and enforce data contracts with producing teams, establish automated validation across completeness, accuracy, consistency, and timeliness, apply confidence scoring to feeds that business and clinical claims depend on, and detect missing, stale, or implausible records upstream of the models.
- Reliability and Operations:Own the operational health of the data platform. Define SLAs and runbooks, instrument monitoring and alerting, and lead incident response and post-incident review.
- Privacy, Compliance, and Lineage:Implement access controls, masking, de-identification, and PHI handling in line with HIPAA and BD policy, partnering with information security, legal, and regulatory on row and column-level security. Maintain lineage, cataloging, and audit trails sufficient for clinical, regulatory, and customer security review, and support change-control and validation requirements for regulated products.
- Technical Leadership: Mentor data engineers and review designs and code. Partner with ML engineers on feature engineering, dataset versioning, and reproducibility, translating operational and clinical context into the features and training datasets models need, and document architecture, schemas, data contracts, and operational runbooks in Confluence.
Required Qualifications:- Bachelor's in Computer Science, Data Engineering, or related field.
- 7+ years of data engineering experience with increasing technical ownership, including end-to-end ownership of production pipelines and data models.
- Demonstrated architectural ownership beyond a single team: setting standards and driving technical decisions that other engineers build on.
- Expert SQL: query authoring, optimization, partitioning, distribution keys, window functions.
- Strong hands-on experience with a cloud data warehouse at production scale, including schema design, performance tuning, workload management, and lake integration. Amazon Redshift (cluster or Serverless, with S3 via Spectrum and COPY/UNLOAD) is our current stack.
- Production experience with AWS data services: S3, Glue, Lambda, Step Functions, Athena, and streaming or CDC via Kinesis, MSK, or DMS.
- Python proficiency for pipeline code, with strong software-engineering habits: tests, code review, modular design, CI.
- Solid grasp of data modeling (Kimball/Inmon, event modeling), governance, lineage, and data quality practice.
- Hands-on feature engineering experience for ML: building, validating, and maintaining feature and training datasets in partnership with ML engineers, including point-in-time correctness and avoiding leakage.
- Ability to learn a business domain deeply enough to translate operational and clinical processes into the right data inputs for modeling, including which signals matter, how they are generated, and where they are unreliable.
- Demonstrated ownership of data quality as an outcome, including validation frameworks, anomaly detection on production feeds, and agreements with upstream producers.
- Practical knowledge of HIPAA and PHI handling, including access control, masking, and de-identification in production environments.
- Operational maturity: SLAs, monitoring, on-call, incident response, and post-incident improvement.
- Working in Agile/Scrum and documenting in Confluence.
Preferred Qualifications:- Master's in Computer Science, Data Engineering, or related field.
- Healthcare or medical-device data experience (HL7/FHIR, device telemetry, operational and clinical data, ERP/SCM feeds).
- Experience in regulated environments with formal change control, computer system validation, audit trail, and data integrity requirements (21 CFR Part 11, GxP, or equivalent).
- Experience with dbt for transformations, and with semantic or metrics layers.
- Experience implementing lakehouse patterns on S3 (Iceberg, Delta, or Hudi).
- Experience supporting time-series and multi-series ML workloads, including temporal feature structures and dataset versioning for reproducibility.
- Familiarity with feature stores or vector stores (Feast, Tecton) for ML use cases.
- Experience building data standardization and entity resolution pipelines using third-party ontologies, NLP, or LLM-assisted mapping (Bedrock or similar).
- Experience with agentic AI pipelines and services, and with building AI-powered workflows and evals.
- Professional use of AI-assisted development tools (Claude Code, Copilot, Cursor) with strong judgment around correctness, security, licensing, and review discipline.
- AWS Data Analytics Specialty certification.
Team Culture:We are building a high-ownership, mission-driven team, energized by the opportunity to solve hard problems that make a difference in patients' lives.
- High ownership - You take full responsibility for outcomes. You identify problems early and drive them to resolution, taking initiative rather than waiting for direction.
- Entrepreneurial spirit - You think and act like an owner: resourceful, creative, willing to challenge assumptions, and consistently focused on impact. You find practical paths forward through difficult problems.
- Mission-driven work ethic - You are motivated by the work we are doing to advance the world of health and improve patient outcomes, and you bring genuine energy and commitment to that mission. You are driven by a real sense of purpose in your work.
- Creative problem-solving - You are comfortable working through ambiguity. You dig in, experiment, iterate, and deliver, and you are energized by the challenge of figuring things out and finding the right solution.
- AI-native mindset -You actively use the latest AI tools to accelerate your own work across coding, research, design, and documentation, and you are eager to keep advancing what is possible with AI in your craft.
Primary Work LocationUSA CA - Irvine Laguna Canyon
Salary Range Information$161,400.00 - $258,200.00 USD Annual