Build our future together:
We are seeking an Associate Director Solutions Partner to be a Scientific Data Product Owner and to lead the holistic data strategy for modelling, connecting, and managing Research and preclinical development scientific data in our enterprise data lake. This data originates from our transactional lab informatics systems (e.g., ELN, LIMS, instrument and analytical platforms, registration systems), and our goal is to transform it from siloed transactional records into well-modelled, connected, and reusable data products.
Sitting at the intersection of science, data engineering, and product management, you will own the vision and roadmap for how scientific data from Research and preclinical development is structured and surfaced within the data lake. You will partner closely with Digital & Technology teams that need connected data products to analyze and report on scientific outcomes, and with data scientists who require curated, connected scientific datasets to power analytics, modelling, and AI initiatives. You will also work hand in hand with data engineering teams to build these connected data products in the lake.
This is a strategic, cross-functional role for someone who understands both the underlying science and the data architecture required to make that science discoverable and actionable at scale, and who is able to productively partner with multiple stakeholders at multiple levels throughout the organization.
When: Based at our Tarrytown, NY location
Where: 4 days onsite per week required
Discover your role:
- Define and drive the holistic data strategy for modelling and managing connected scientific data in the data lake, ensuring data flowing in from transactional lab informatics systems is harmonized, contextualized, and connected.
- Partner with data engineering to define and shape the ingestion, transformation, and integration pipelines (ETL/ELT) that move data from transactional lab informatics systems into the data lake and turn it into connected, production-grade data products.
- Own the product vision, roadmap, and backlog for scientific data products, prioritizing based on scientific value, downstream demand, and organizational impact.
- Design and govern data models that capture the relationships across scientific entities (e.g., compounds, biologics, samples, assays, experiments, results, instruments, and study metadata) so that data is connected rather than isolated.
- Translate the needs of scientific, Digital & Technology, and analytics stakeholders into clear data product requirements, acceptance criteria, and delivery plans, and drive outcomes that are aligned with all stakeholders, as well as with the corporate Digital and Transformation strategy.
- Partner with Digital & Technology teams that consume connected data products from the data lake for analysis and reporting, ensuring data products meet their integration, quality, and access requirements.
- Partner with data scientists seeking connected scientific datasets, ensuring datasets are analytics-ready, well-documented, and fit for modelling and AI/ML use cases.
- Establish and uphold data quality, governance, lineage, metadata, and FAIR (Findable, Accessible, Interoperable, Reusable) principles across scientific data products.
- Conceive, elicit, and champion the use of AI-driven approaches for searching and discovering structured scientific data, improving how users find and connect relevant datasets.
- Act as the primary point of contact and advocate for scientific data products, gathering feedback and continuously improving usability, coverage, and value.
This role requires:
- BS/BA degree and/or MS degree in a related field required.
- 10+ years of progressive experience managing scientific laboratory data, ideally within the biopharmaceutical or life sciences industry.
- Strong scientific background (e.g., chemistry, biology, molecular biology, biochemistry, immunology, pharmacology, or a related discipline), with the ability to understand the meaning and context of the data being modelled.
- Hands-on experience with data modelling and an understanding of how to connect data across multiple source systems.
- Ability to drive diverse stakeholders to alignment on desired outcomes, and to influence others at multiple levels without direct authority.
- Working knowledge of data engineering concepts — data pipelines, ETL/ELT, and data transformation — sufficient to define requirements for and collaborate effectively with data engineers.
- Strong SQL skills, plus proficiency in a primary Databricks language (e.g., SQL, Python/PySpark).
- Familiarity with transactional lab informatics systems such as ELN, LIMS, and instrument or analytical data platforms.
- Product ownership or product management experience, including roadmap definition, backlog prioritization, and stakeholder management.
- Strong communication skills and the ability to bridge scientific, technical, and analytics audiences.
- Advanced degree in a biology discipline required; PhD in molecular biology, biochemistry, genetics, or immunology preferred.
Strongly Desired
- Hands-on experience with Databricks (or a comparable lakehouse platform) for managing and delivering data products.
- Experience leveraging AI to search, discover, and interrogate structured data.
- Experience with knowledge graphs, ontologies, controlled vocabularies, or semantic data models for connecting scientific entities.
- Understanding of data lake / lakehouse architectures and modern data engineering practices.
- Experience working with data scientists and analytics teams to deliver analytics-ready datasets.
Nice to Have
- Familiarity with data governance frameworks and FAIR data principles in regulated and unregulated environments.
- Experience with metadata management, data cataloging, and data lineage tooling.
- Experience with AI/ML workflows and the data requirements that support them.
- Awareness of regulatory and compliance considerations relevant to pharmaceutical R&D data.
Salary Range (annually)
$216,100.00 - $360,200.00