About the roleWe're looking for a Data Scientist to help build the machine learning capabilities that will power the next phase of our R&D platform. In this role, you'll report to the Director of R&D and work side-by-side with laboratory scientists, applying machine learning and chemoinformatics to turn experimental data into testable predictions that accelerate scientific discovery. You'll have the opportunity to own models from initial concept through deployment, influence the technical direction of our ML infrastructure, and build scalable data pipelines that become foundational to the team's work.
If you're excited by applying cutting-edge machine learning to real-world scientific problems in a highly collaborative, interdisciplinary environment, this is a chance to have a meaningful impact on both our research strategy and the company's future growth.
This is a full-time, on-site role based out of Berkeley, CA.
What You'll Do- Design, build, and deploy machine learning models to predict chemical properties and support de novo design of small molecules and peptides, taking models from concept through validation and production.
- Own the development and maintenance of scalable data pipelines that ingest, curate, and prepare experimental screening data for machine learning applications, ensuring reproducibility and reliability.
- Partner closely with laboratory scientists to translate experimental questions into machine learning approaches, generate actionable predictions, and prioritize new compounds or target areas for validation.
- Develop and improve the team's machine learning infrastructure, including model training, evaluation, retraining, monitoring, and continuous improvement as new experimental data becomes available.
- Apply chemoinformatics tools and molecular representations to engineer features, evaluate model performance, and improve predictive accuracy across R&D workflows.
- Communicate insights and recommendations through clear visualizations and presentations, helping both technical and non-technical stakeholders understand model performance and scientific findings.
- Contribute to the technical direction of the machine learning platform, collaborating with scientists and other technical team members to identify opportunities for new models, datasets, and workflows that accelerate research.
What You'll Bring- Ph.D. or Master's in Computational Chemistry, Chemoinformatics, or a related field and 2-4 years of post-graduate experience, with a strong foundation in chemical structure representation and molecular property prediction.
- Proven track record in building and deploying machine learning (ML) and deep learning models to predict chemical characteristics and drive de novo design for chemical compounds with superior properties. Experience working with small molecule or peptide datasets.
- High proficiency with using chemoinformatics toolkits (e.g. RDKit) and molecular descriptors, fingerprints, and structural data curation. Experience with structure or ligand-based tools (e.g. REINVENT) and molecular dynamics simulations is a plus
- Fluent in Python and data science stacks, such as numpy, pandas, scipy. Hands on experience with SQL and managing data pipelines to ensure models are reproducible and scalable.
- Strong communication and interpersonal skills. Able to maintain highly productive working relationships with laboratory scientists to ingest screening data, refine models, and recommend testable predictions.
- Able to work independently to lead efforts and turn conceptual goals into structured data pipelines.
- Experience with data visualization tools to communicate model results and predictions to technical and non-technical stakeholders.
Essential traits:
- Highly collaborative & comfortable with interdisciplinary communication: Success in this role depends on working closely with laboratory scientists and other technical partners to translate scientific questions into machine learning solutions and communicate results that drive research decisions.
- Self-directed and results-oriented problem-solver: You'll be expected to independently identify opportunities, overcome technical challenges, and move projects from concept to implementation while delivering meaningful outcomes in a fast-paced startup environment.
- Able to own and navigate complex and ambiguous tasks: Many of the problems you'll tackle won't have established solutions, requiring you to bring structure to uncertainty, make sound technical decisions, and adapt as new data and discoveries emerge.
Compensation Aralez Bio's salary range for this position is $165,000 to $180,000 per year, based on your performance and experience. In addition, your total rewards package will include equity and benefits.
Benefits at Aralez Bio- Medical / dental / vision insurance coverage
- 401k
- Flexible Spending Account
- Paid time off
- Weekly catered lunch
- Professional Development opportunities
- Annual company retreat