About the RoleOpenAI's Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models.
We're looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience.
You'll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world's largest AI compute environments.
Key Responsibilities- Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency.
- Develop forecasting models for inference demand across products, regions, and model families.
- Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities.
- Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies.
- Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs.
- Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions.
- Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps.
- Communicate technical findings clearly to both engineering teams and executive leadership.
Qualifications- MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience).
- 5+ years of experience working in the infrastructure data science space.
- Strong expertise in Python and SQL.
- Experience building forecasting, optimization, or predictive models.
- Strong understanding of experimentation, statistical inference, and causal analysis.
- Experience communicating analytical insights to executive stakeholders.
Preferred Skills- Capacity planning
- Distributed systems
- AI infrastructure
- Datacenter design and buildout
- Queueing theory
- Time-series forecasting
- Operations research
- Supply-demand modeling
- Reinforcement learning for resource allocation
- Cost optimization