ML Platform powers hundreds of use cases and billions of inferences per day across Discovery, Safety, Economy, and Creation. We build the primitives that let teams train, evaluate, deploy, and operate models quickly and safely-so a new ML idea can reach production in weeks or less. We9re looking for a Principal Platform Engineer who treats platform as a product: someone who can turn complex ML/AI infrastructure into clear, durable APIs, SDKs and easy-to-use CLIs/UIs that our internal developers love.
This role blends product thinking, developer experience, backend engineering, and infrastructure at scale. Hands-on ML experience is a plus; a track record of building internal platforms that developers love is a must.
You Are- Have 7+ years of professional experience and have a wealth of system design experience upon which to draw to build a scalable, reliable ML platform for all of Roblox.
- A Code Machine; you love not only to design and communicate ideas but also to actually ship product. You obsess about user feedback, and constantly drive towards getting platform features in customers hands.
- Proficient in API design and developer experience-gRPC/REST APIs, SDKs, CLIs, and simple UIs that developers love to use.
- Good to have - experience with the end-to-end ML model lifecycle such as model serving, training, model CI/CD, and GPU resources management, and have built ML platform features that are delightful to use.
- You9re passionate about infrastructure-as-code and automating painful manual processes. You push for platform solutions that don9t leave on-call teams carrying the burden of design choices.
- Passionate about supporting ML engineers to meet and understand their needs, and translating them into clean, durable platform abstractions.
- A clear communicator who excels at both written and verbal communication across differing levels of technical detail.
- Bachelor9s degree in Computer Science, Computer Engineering, Data Science, or a similar technical field.
You Will- Own platform as a product and set direction end to end: Define requirements, write RFDs, and ship APIs, SDKs, CLIs, and UIs that make ML[redacted] easy to adopt.
- Bootstrap and maintain core ML Platform components: Serving Layer, Model Registry, Pipeline Orchestrator, and Training/Inference control planes.
- Set technical strategy and oversee development of high scale and reliable infrastructure systems, with clear SLOs for latency, availability, and cost.
- Design great developer experiences with paved-road templates, golden paths, opinionated defaults, and clear docs to reduce time-to-first-production.
- Instrument the platform to measure adoption, friction, reliability, and cost; use data to prioritize roadmap and validate outcomes.
- Partner across organizations (ML Engineering, Data Science, Infra/SRE, Security, Finance) to optimize performance, safety, and spend, especially for GPU-intensive training and high-QPS inference.
- Propose and implement new platform tooling to improve time to production for MLEs across the full ML lifecycle.
- Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices.
- Mentor junior and senior engineers, lead design reviews, and drive cross-team architectural decisions that last.
For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range. This pay range is subject to change and may be modified in the future. All full-time employees are also eligible for equity compensation and for benefits as described on
this page.
Annual Salary Range
$260,300-$345,040 USD
Roles that are based in an office are onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday (unless otherwise noted).