The OpportunityFirefly Foundry is Adobe's enterprise managed-service offering for custom multimedia generative AI - deep-tuned image, video, and 3D models built on each customer's IP, paired with creative production workflows and a media-intelligence layer, and deployed across new and existing Adobe surfaces and products, including Firefly, Photoshop, Illustrator, Express, Stock, and Premiere.
We are hiring a Principal Machine Learning Engineer to serve as the technical lead for our GenAI Services area. This is not a model-training or research role - it is the senior-most hands-on engineering authority over how our generative models are architected, optimized, and served at enterprise scale. You will set the inference architecture and technical standards that a growing organization of engineers builds against, co-develop and optimize the inference code that makes those systems fast and cost-efficient, and architect the APIs and product backend that let Adobe's first-party and third-party models reach both internal applications and external plugin integrations. Where the Director owns the multi-year technical strategy, headcount, and company roadmap for the org, you own the architecture, technical depth, and hands-on execution that make that strategy real - spanning multiple engineering teams without owning their people management.
What this role owns - The technical architecture for composing, optimizing, and serving heterogeneous generative model pipelines - LLMs, diffusion and transformer-based image/video models, RAG and retrieval systems, multi-turn agentic flows, and 3D/mesh pipelines - across the GenAI Services area.
- The optimization strategy for inference performance: latency, throughput, and cost-to-serve across model families and GPU fleets.
- The system design standards for pipeline composition, multi-tenant serving, and the product backend/API and plugin surface that integrates first-party and third-party generative models into Adobe's flagship products.
- Technical direction across multiple engineering teams as the principal authority on architecture and design - a cross-team scope, distinct from the Director's org-wide roadmap and management ownership.
Who you will partner with - Applied Science - to translate research models and emerging techniques into production-grade inference architecture.
- Director, ML Engineering and ML Engineering leadership - to align technical architecture with organizational strategy and priorities.
- Product Managers and TPMs - to define and deliver against the roadmap for GenAI services and APIs.
- Firefly Foundry Studio and AI Platform - to translate creative production workflows into performant services and to align on shared infrastructure and serving primitives.
What you will do - Lead the development of core GenAI services and APIs that integrate a wide range of first-party and third-party generative models into Adobe's flagship products.
- Architect ML serving workflows for enterprise-scale model customization, deployment, and ecosystem integration - including externalizable, self-serve fine-tuning flows.
- Co-develop and optimize GPU-accelerated inference pipelines - prioritizing latency, throughput, scalability, and reliability - using tools such as PyTorch, CUDA, Triton, and TensorRT.
- Design and architect the product backend and plugin ecosystem that lets internal applications and external integrations consume Firefly Foundry's model services.
- Provide hands-on technical leadership: guide engineers through architecture, design, implementation, and best practices, and mentor a growing organization of ML engineers.
- Research and evaluate emerging inference and MLOps technologies - serving runtimes, quantization, GPU scheduling - to improve engineering velocity and system performance.
- Lead design reviews and set technical standards, ensuring high reliability and maintainability across systems.
- Drive cross-functional alignment with Product Managers, TPMs, and engineering leaders to define and deliver on the roadmap.
- Foster a culture of technical excellence and continuous improvement across the organization.
What you bring - MS or PhD in Computer Science, Machine Learning, or a related field - or equivalent industry experience.
- 8+ years of experience in machine learning engineering, including production-scale deployment and serving - not training or research experimentation.
- 3+ years leading the technical direction of large-scale, GPU-intensive GenAI inference systems - serving, architecture, and optimization.
- Deep experience with inference frameworks and tools such as PyTorch, CUDA, Triton, TensorRT, Nvidia Dynamo, and Python.
- Strong understanding of generative model architectures - diffusion models, transformers, GANs, LLMs - sufficient to make architecture and optimization calls and reason about output quality, in partnership with Applied Science.
- Proven experience architecting multi-model pipelines and serving them behind APIs at enterprise scale.
- Experience designing product backend systems and plugin architectures consumed by internal applications and external integrations.
- Proven success leading cross-functional teams through complex, high-stakes technical initiatives, with a track record of driving alignment in matrixed organizations.
- Excellent communication and technical leadership skills.
Preferred Qualifications - Experience with model serving, orchestration, and GPU resource management in large-scale environments.
- Hands-on expertise in Kubernetes, distributed systems, and MLOps platforms.
- Experience with RAG architectures and multi-turn, agentic conversational systems.
- Experience with quantization, distillation, or other model-optimization techniques for inference.
Education - Master's or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience building and leading production-scale ML systems.
#FireflyGenAI