About the RoleA Manager Data Science is an emerging subject matter expert in their domain. They lead a team of junior members to support their development and work product. They are mindful of best practices and train their team in the execution of those best practices. They manage a team to define new best practices and innovative approaches to new business problems or use cases.
Responsibilities
- Lead the design and execution of LLM training and fine-tuning projects, including model selection, training strategy, experimentation, and evaluation.
- Oversee the preparation of high-quality training datasets, including data collection, cleaning, deduplication, annotation, and quality validation.
- Develop and optimize supervised fine-tuning and parameter-efficient fine-tuning workflows; apply preference optimization methods where appropriate.
- Establish evaluation frameworks to assess factual accuracy, instruction following, domain relevance, safety, and performance on business-specific tasks.
- Diagnose training issues and improve model quality, training stability, GPU utilization, and computational efficiency.
- Manage and mentor data scientists, review technical work, and establish reproducible development practices.
- Partner with product, engineering, and domain experts to define requirements and support model deployment and monitoring.
- Manage project priorities, timelines, and compute resources, and communicate results and tradeoffs to stakeholders.
Requirements- LLM fundamentals: Strong understanding of transformer architectures, attention mechanisms, tokenization, language modeling objectives, and the differences between pretraining, continued pretraining, and fine-tuning.
- Programming and frameworks: Strong Python and PyTorch skills, with practical experience using Hugging Face Transformers, Datasets, or equivalent tools.
- Hands-on LLM training: Demonstrated ability to implement supervised fine-tuning (SFT), configure training objectives and loss masking, tune hyperparameters, and select model checkpoints.
- Efficient fine-tuning: Practical experience with parameter-efficient fine-tuning (PEFT), including LoRA or QLoRA, and an understanding of their quality, memory, and compute tradeoffs.
- Training data engineering: Ability to build instruction-response datasets, apply chat templates, manage sequence lengths and packing, and prevent data leakage and evaluation contamination.
- GPU and distributed training: Experience training models across multiple GPUs using frameworks such as PyTorch FSDP or DeepSpeed, including mixed precision, gradient accumulation, and gradient checkpointing.
- Evaluation and debugging: Ability to design reliable benchmarks and human evaluations, analyze model errors, and troubleshoot unstable loss, overfitting, and GPU memory issues.
- Reproducibility: Experience with experiment tracking, dataset and model versioning, checkpoint management, and documented training pipelines.
Work in a Way That Works for YouWe promote a healthy work/life balance across the organisation. We offer an appealing working prospect for our people. With numerous wellbeing initiatives, shared parental leave, study assistance and sabbaticals, we will help you meet your immediate responsibilities and your long-term goals.
Working PatternWorking flexible hours - flexing the times when you work in the day to help you fit everything in and work when you are the most product
ive.