What’s the role?
We are looking for a Senior Software Engineer, Machine Learning to help build and operate the infrastructure that trains, hosts, and serves Etsy's machine learning models — including our internal serving platform, and a growing portfolio of predictive and generative (open source LLM) models.
You won't just be deploying models; you'll be architecting the systems that make model serving fast, reliable, and scalable for millions of inferences per second. Your work directly enables Applied Scientists and Engineers across Etsy to train, host, and iterate on models with confidence, from classic ML predictors to large open-source language models.
This is a full-time position in ML Enablement team, reporting to its Engineering Manager.
What’s this team like at Etsy?
- The Machine Learning Enablement initiative builds the core infrastructure that turns complex ML workflows into seamless, self-service platforms for Etsy’s Applied Scientists and Engineers.
- The Models Training and Serving Platform team builds and operates the core infrastructure that powers model training and serving across Etsy, including Barista, our internal model-serving platform, along with a range of hosted modelsets and open-source LLMs.
- We own the path from a trained model to a production-ready, observable, and scalable serving endpoint — spanning infrastructure, deployment tooling, and reliability.
- We work on significant, complex challenges that intersect multiple critical teams and systems, where you can make a rewarding impact.
- We are light on process and heavy on collaboration, working with many partner teams within our org and beyond in order to improve our leverage.
- We are a platform team with a product driven mentality, driving innovation using our Machine Learning systems in effective and creative ways.
What does the day-to-day look like?
- Write high-quality, production-grade Python code (additional languages a plus); participate in code reviews and pair programming; contribute to and help drive architectural decisions.
- Build, operate, and improve infrastructure for training and serving ML models on GCP and Kubernetes, with a focus on scalability, reliability, and observability.
- Design and maintain serving infrastructure for open-source and internally hosted LLMs, including model deployment, resource management, and performance tuning.
- Apply working knowledge of ML fundamentals — including neural network deep learning as well as latest transformer architectures along with prediction and inference systems — to make sound infrastructure and design decisions.
- Partner cross-functionally with Applied Scientists to understand model training and serving needs, and translate that understanding into infrastructure that removes friction from their workflows.
- Contribute to system design discussions, weighing tradeoffs across performance, cost, and reliability for infrastructure serving 100M+ users.
- Thoughtfully use generative AI and other productivity tools to work efficiently, with a focus on learning and intentional contribution.
- Of course, this is just a sample of the kinds of work this role will require! You should assume that your role will encompass other tasks, too, and that your job duties and responsibilities may change from time to time at Etsy's discretion, or otherwise applicable with local law.
Qualities that will help you thrive in this role are:
- Bachelor's degree in Computer Science, Applied Statistics, Mathematics, Electrical Engineering, or a related quantitative field, or equivalent professional experience.
- 5+ years of professional experience building, iterating on, and troubleshooting complex backend and infrastructure systems.
- Strong software engineering fundamentals, including solid command of algorithms and data structures, with the ability to write production-ready code in Python.
- Hands-on experience with cloud infrastructure (Google Cloud preferred) and Kubernetes, including deploying and operating production workloads.
- Familiarity with observability tooling (metrics, logging, tracing) for monitoring and debugging distributed systems.
- Working knowledge of machine learning fundamentals and concepts, with awareness of LLM serving infrastructure and hosting open-source models.
- Basic understanding of transformer architectures, predictors, and how model design choices affect serving infrastructure.
- Comfort with system design for large-scale, high-availability infrastructure.
Compensation & Benefits
In addition to salary, you'll be eligible for an equity package, an annual performance bonus, and our that support you and your family. Base salary is determined by your location, skills, and experience.
Etsy uses two geographic pay zones to reflect the labor markets where we compete for talent: one called Geo-1, covering the New York City metropolitan area, San Francisco Bay Area, and Seattle metropolitan area, and one called Geo-2, covering all other U.S. locations. Salary ranges for this role vary by zone and are listed below.
Location
We prioritize candidates based within commutable distance of Etsy’s Brooklyn Office. Depending on proximity, we ask team members to come in once or twice per week. Remote candidates outside commutable distance may be considered on a case-by-case basis, with the expectation of occasional travel to the office. Learn more details about our work modes and workplace safety policies .
Salary Range:
• Geo-1: 182,000.00 - 246,000.00 USD Annual
• Geo-2: 164,000.00 - 222,000.00 USD Annual
What's Next
If you're interested in joining the team at Etsy, please share your resume with us and feel free to include a cover letter if you'd like. As we hope you've seen already, Etsy is a place that values individuality and variety. We don't want you to be like everyone else -- we want you to be like you! So tell us what you're all about.