Role Overview:This role focuses on the curation, validation, and management of chemical structure data, alongside the design and development of Python-based applications and automated workflows for scientific data processes. The successful candidate will apply cheminformatics expertise to support molecular analysis, data quality, and AI/ML initiatives within a research and development environment.
Key Responsibilities:- Curate, validate, and maintain chemical structure data across research platforms and scientific workflows.
- Perform structure standardization and normalization, including salt/solvent handling, valence checks, charge normalization, tautomer treatment, stereochemistry review, and duplicate identification.
- Support compound registration, structure correction, modification, and molecular metadata management.
- Investigate and resolve data-quality issues in collaboration with medicinal chemists, computational scientists, and informatics teams.
- Define and apply repeatable quality-control rules for chemistry datasets used in analytics and AI/ML initiatives.
- Design and develop Python-based applications, APIs, libraries, and automated workflows for chemistry and scientific data processes.
- Build scalable pipelines for structure ingestion, transformation, validation, standardization, and export.
- Integrate chemistry platforms, scientific databases, and enterprise or cloud services using APIs and data-integration patterns.
- Write modular, tested, maintainable code and contribute to code reviews, documentation, release practices, and production support.
- Create scientist-friendly utilities, notebooks, or lightweight interfaces to reduce manual curation effort and improve traceability.
- Use cheminformatics toolkits such as RDKit, ChemAxon, or equivalent technologies for molecular analysis and structure processing.
- Develop capabilities for substructure and similarity search, scaffold analysis, descriptor calculation, compound clustering, and molecular property computation.
- Apply chemistry knowledge to ensure software outputs are scientifically meaningful and fit for use.
- Contribute to data preparation and feature-generation workflows supporting predictive modeling, molecular design, and AI-enabled research.
- Partner with chemists, informaticians, data scientists, product owners, and software engineers to translate scientific needs into technical solutions.
- Participate in solution design, estimation, development, testing, validation, deployment, and operational support.
- Communicate technical decisions, chemistry-data risks, and delivery status clearly to scientific and technology stakeholders.
- Promote reusable engineering patterns, chemistry-data standards, and continuous improvement.
Required Skills:- Strong professional experience developing production-quality solutions in Python.
- Solid grounding in organic chemistry and chemical structure representation, including stereochemistry, salts, tautomers, charges, and structure-quality concepts.
- Hands-on experience with chemistry curation, compound registration, structure standardization, or molecular data management.
- Experience with Python scientific and data libraries such as pandas and Jupyter, plus SQL and REST APIs.
- Experience with at least one cheminformatics toolkit or chemistry platform, such as RDKit, ChemAxon, BIOVIA Pipeline Pilot, Dotmatics, Signals, or an equivalent platform.
- Ability to work across scientific and technical teams and explain complex concepts to varied audiences.
Qualifications:- Master's degree or PhD in Chemistry, Cheminformatics, Computational Chemistry, Pharmaceutical Sciences, Computer Science with substantial chemistry experience, or a related discipline. Equivalent relevant industry experience may be considered.
Preferred Skills:- Experience supporting pharmaceutical or life-sciences R&D, particularly small-molecule discovery or chemical synthesis workflows.
- Familiarity with compound registration systems, ELN platforms, chemical inventory tools, or scientific data platforms.
- Knowledge of cloud platforms, containers, CI/CD, automated testing, and secure software-development practices.
- Exposure to AI/ML applications involving molecular data, property prediction, molecule generation, or synthesis planning.
- Understanding of scientific data governance, auditability, validation, and regulated-environment expectations.