Scientific Data Engineer

LBL$131K — $161K *
Pharmaceuticals & Biotech
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Minimum 5 years experience with a Bachelor's degree in a relevant field or 3 years with a Master's; PhD is also acceptable.
  • Experience in scientific software development within a research context.
  • Hands-on experience in software production environments, especially with scientific data.
  • Strong programming skills in Python; familiarity with C++ or Javascript is a plus.
  • Experience in contributing to open source projects and working with large codebases.
  • Proficient with Git and CI systems, such as GitHub or GitLab.
  • Strong communication skills and ability to collaborate with non-computing domain experts.

Responsibilities

  • Design and implement user-friendly software for data management and analysis.
  • Collaborate with experts to develop FAIR data models for bioscience applications.
  • Structure and document biological data for machine learning and AI applications.
  • Engage with Neurodata Without Borders and LinkML development communities.
  • Maintain and manage open source software products and coordinate development priorities.
  • Design cloud solutions for visualizing and analyzing complex biological datasets.
  • Train scientists and engineers on developed software during events.

Benefits

  • Exceptional health and retirement benefits including pension or 401K options.
  • Inclusive culture that emphasizes belonging and teamwork.
  • Accrual of vacation and sick leave plus an annual Winter Holiday Shutdown.
  • Parental bonding leave for all parents, regardless of gender.
  • Pet insurance available for employees.
Full Job Description
The Computational Biosciences Group has an immediate opening for a software and data engineer in the area of multi-modal data modeling and analysis with applications to bioscience research. You will develop new methods and software tools that enable scientific knowledge discovery using modern data management and machine learning technologies and advance the state-of-the-art in data-intensive analysis. Your projects will focus on the domains of omics/structural biology data and neurophysiology data. Under limited instruction, you will be part of an experienced team conducting R&D in the areas of FAIR data science, AI, and modern methods for data understanding. You will be working as part of a multi-disciplinary team composed of computer scientists, data scientists, and bioinformaticians. Please note this is a scientific software/data engineering position- it is not a pure machine learning or AI research position, and it is not a pure data science or analytics position.

You will:
  • Design and develop user-friendly software packages for scientific data management and analysis
  • Work with domain experts to develop FAIR data models (i.e., models of the structure organization of the data) and management solutions for bioscience applications
  • Support machine learning and AI use of biological data by making it well-structured, documented, and efficiently accessible
  • Work closely with the community of developers of the Neurodata Without Borders and LinkML open source data ecosystems, as well as the Joint Genome Institute.
  • Maintain and manage open source software products, including managing development priorities, software releases, continuous integration, and testing
  • Design, implement and maintain high performance computing and cloud solutions for visualization and analysis of complex biological data
  • Develop machine learning and AI solutions for analysis of biological data in close collaboration with diverse teams of scientists
  • Train scientists and research software engineers in the use of the developed software products at workshops and conferences
  • Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
  • Network with senior internal and external personnel in their own area of expertise.


We are looking for:
  • Typically requires a minimum of 5 years of related experience with a Bachelor's degree in computer science, data science, machine learning, bioinformatics, or equivalent; or 3 years and a Master's degree; or equivalent work experience designing and developing software for data modeling or analysis; or a PhD in a relevant STEM field
  • Demonstrated experience developing software in a scientific or research context, such as in a research group, a scientific user facility, or on a scientific software project
  • Demonstrated hands-on experience in a production environment, developing scientific software, scientific data models, or scientific data pipelines
  • Strong programming experience in Python. Working proficiency in C++ or Javascript is a plus.
  • Experience testing large code bases
  • Experience contributing to community-driven open source software
  • Demonstrated experience in one or more of the following areas: data management, scientific data analysis, machine learning
  • Works well in a collaborative team environment
  • Demonstrated capability with the Git version control and continuous integration systems, such as GitHub or GitLab
  • Ability to work effectively with domain scientists whose expertise is outside computing, and to translate their requirements into technical designs.
  • Excellent oral and written communication skills.
  • Demonstrated ability to work effectively as part of a cross-disciplinary team.


Desired skills/knowledge:
  • Master's or PhD in Computer Science or related field, with 5 or more years of professional experience designing and developing scientific data modeling or analysis software
  • Experience working with modern scientific data formats and database systems, such as HDF5, Zarr, MongoDB, PostgreSQL, MySQL, and Redis
  • Experience with Neurodata Without Borders, LinkML, or similar software ecosystems
  • Experience working with large biological data, such as in the areas of neurophysiology, microbiology, genomics, or protein design
  • Experience designing or working with structured data models, schemas, ontologies, or data standards
  • Familiarity with FAIR data principles, persistent identifiers, provenance, and controlled vocabularies and ontologies
  • Experience preparing scientific datasets for use by machine learning pipelines or LLM-based agents
  • Experience working with cloud object storage, cloud computing, High-Performance Computing, data lakehouse architecture, or containerization.
  • Experience developing web-based graphical user interfaces (GUIs) or application programming interfaces (APIs) for scientific data analysis and management


How to apply:

In addition to your resume, please include a brief cover letter (max 300 words) that describes one scientific software project you have contributed to, or one scientific data model, schema, standard, or pipeline you have helped design or implement. Describe your specific role and what the technical challenge was. Thesis, internship, research lab, and personal projects count. If the code is public, include a link.

We're here for the same mission, to bring science solutions to the world. Join our team and YOU will play a supporting role in our goal to address global challenges! Have a high level of impact and work for an organization associated with 17 Nobel Prizes!

Why join Berkeley Lab?

We invest in our employees by offering a total rewards package you can count on:
  • Exceptional health and retirement benefits, including pension or 401K-style plans
  • A culture where you'll belong - we are invested in our teams!
  • In addition to accruing vacation and sick time, we also have a Winter Holiday Shutdown every year.
  • Parental bonding leave (for both mothers and fathers)
  • Pet insurance


Additional information:
  • Appointment type: This is a full-time, 2 years, term appointment with the possibility of extension or conversion to Career appointment based upon satisfactory job performance, continuing availability of funds and ongoing operational needs.
  • Salary range: The expected salary for this position is $131,760 - $161,064, which fits into the full salary of $117,132 - $197,676 depending upon the candidate's skills, knowledge, and abilities. This includes education, certifications, and years of experience.
  • Work modality: Work may be performed on-site, or hybrid. The primary location for this role is Lawrence Berkeley National Lab, 1 Cyclotron Road, Berkeley, CA. Work must be performed within the United States. A REAL ID or other acceptable form of identification is required to access Berkeley Lab sites (for more information click here).


Want to learn more about working at Berkeley Lab? Please visit: careers.lbl.gov

About LBL

LBL Careers

Joining LBL offers an unparalleled opportunity to become part of a leading team of professionals dedicated to pioneering innovation and digital transformation. LBL stands as a beacon of excellence, offering a range of job opportunities that cater to various skills and career aspirations.

Explore Career Opportunities

LBL’s dynamic career paths empower professionals to navigate their professional growth with confidence. Whether through full-time positions, internships, or leadership roles, LBL is committed to fostering a culture of growth and learning.

Innovation and Professional Growth

At LBL, innovation isn’t just a buzzword; it's the cornerstone of their mission. The company encourages its team to push the boundaries of technology and strategy, ensuring that every member has the opportunity to contribute to groundbreaking projects.

Diversity and Inclusion

Diversity training and inclusion are at the heart of LBL’s employment strategy. The company believes that a diverse team is a strong team, and actively works to create an environment where all voices are heard and valued.

Benefits and Culture

LBL is renowned for its vibrant culture and comprehensive benefits package designed to support the team in all aspects of life—both professional and personal. From health benefits to flexible work policies, LBL ensures that the team not only excels at work but also enjoys a balanced life.

Networking and Development

Career advancement at LBL is fueled by robust professional networking and development programs. These initiatives are tailored to hone skills, enhance leadership capabilities, and ensure that every team member can achieve their career goals.

Join the LBL Team

LBL is actively hiring and looking for individuals who are passionate, curious, and driven. Explore the open positions that match your skills and interests. Engage with a company that values innovation and offers the tools needed to succeed in a competitive market.

Stay Connected with LBL Jobs

Stay informed about the latest in career opportunities and industry trends by subscribing to LBL job alerts. Tailor your preferences to receive updates that align with your professional interests and career goals.

Prepare for Your Interview

Aspiring candidates can look forward to a transparent interview process that assesses a range of competencies from technical skills to creative thinking. Ensure your resume highlights relevant experiences and skills to stand out in the LBL hiring process.

Career Insights and Tips

Gain insights from industry leaders and get ahead with career tips directly from the professionals at LBL. These resources are invaluable for those looking to make a significant impact in their professional journey.

Explore LBL Careers Today

Discover the exciting and rewarding career opportunities at LBL. Whether you’re seeking an internship or a managerial position, LBL offers a path for everyone. Join a team that’s dedicated to leadership, professional growth, and innovation in the digital era.
Learn more about LBL

Similar Jobs

More Jobs at LBL

More Pharmaceuticals & Biotech Jobs

Find similar Scientific Data Engineer jobs: