Design, develop, and productionize scalable bioinformatics workflows supporting pangenome graph construction, haplotype expansion, population-scale imputation, and genomic data quality control. This role focuses on transforming research-grade analyses into robust, reusable, and cloud-enabled pipelines that support large-scale multi-crop genomics programs.
Requirements- Develop and maintain production-grade bioinformatics workflows for pangenome, haplotype, imputation, and QC processes.
- Convert research scripts and manual analyses into automated, version-controlled, and reproducible pipelines.
- Work with genomic data formats including FASTA, GFF/GTF, VCF, BAM/CRAM, haplotype outputs, and associated metadata.
- Implement workflows using cloud-native AWS services, leveraging S3 storage and scalable batch execution.
- Build validation, logging, provenance tracking, and error-handling capabilities into workflows.
- Collaborate with scientists and domain experts to ensure biological accuracy and usability of outputs.
Required Skills- Strong programming experience in Python, Nextflow, Snakemake, WDL, or similar workflow frameworks.
- Experience processing large-scale genomics datasets and common bioinformatics file formats.
- Familiarity with AWS cloud services and distributed computing environments.
- Understanding of software engineering best practices, including Git, testing, CI/CD, and documentation.
Key Deliverables- Production-ready workflow modules and reusable pipeline templates.
- Automated QC and validation reports.
- Execution and operational documentation.
- Reliable mechanisms for workflow monitoring, reprocessing, and onboarding of new datasets.
BenefitsWe offer a competitive compensation package including accrued vacation, medical, dental, vision, 401k with company matching, life insurance, and flexible spending accounts.
If you share these sentiments and are prepared for the atypical, then Zifo is your calling!