OpenAI

Software Engineer - Data Aquisition (systems)

OpenAI$130K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of software engineering experience with a focus on infrastructure and systems engineering.
  • Strong understanding of systems fundamentals and scaling infrastructure in practice.
  • Proficiency in Linux, networking, Kubernetes, and provisioned distributed systems operations.
  • Ability to write software for automation and improvement of infrastructure systems.
  • Independent and pragmatic approach to ambiguous infrastructure work.

Responsibilities

  • Build and operate infrastructure for research workloads and services.
  • Support and enhance data infrastructure, processing, and observability systems.
  • Improve cluster bootstrapping and deployment workflows.
  • Debug complex issues across networking, compute, and storage layers.
  • Develop software and automation for improved operational reliability.
  • Collaborate with researchers and engineers to meet system requirements.
  • Drive ownership of critical systems from problem definition to execution.

Benefits

  • Flexible working environment with opportunities for independent work.
  • Engagement with cutting-edge research and novel technologies.
  • Collaboration with diverse teams, including researchers and engineers.
  • Focus on personal growth and development in a high-ownership culture.
Full Job Description
About the Role

As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time.

This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work.

We expect you to:
  • Build and operate reliable infrastructure for research workloads and research-facing services.
  • Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services.
  • Improve cluster bootstrapping, provisioning, automation, and deployment workflows.
  • Debug issues across networking, compute, storage, orchestration, and service reliability layers.
  • Build software and automation that reduce manual operational work and improve system reliability.
  • Partner closely with researchers, infrastructure engineers, and service owners to understand system needs and translate them into durable solutions.
  • Help evolve existing infrastructure toward more scalable, maintainable, and standard patterns.
  • Take ownership of critical systems and drive work independently from problem definition through execution.


You might thrive in this role if you:
  • Have strong systems fundamentals and understand how infrastructure scales in practice.
  • Are comfortable with Linux, networking, Kubernetes, provisioning, and distributed systems operations.
  • Can write software to automate, debug, and improve infrastructure systems.
  • Have a strong execution mindset and can independently drive ambiguous infrastructure work.
  • Enjoy supporting a wide surface area of systems, from research tooling to platform services.
  • Are pragmatic about when to build custom systems versus using existing, well-supported tools.
  • Care about building reliable systems that make researchers faster and reduce operational friction.


Nice to have:
  • Experience with PXE boot, cluster provisioning, bare-metal infrastructure, or large-scale fleet management.
  • Experience operating Kubernetes or similar orchestration systems at scale.
  • Experience with infrastructure-as-code, CI/CD, observability, or deployment automation.
  • Experience supporting search infrastructure, data platforms, ingest systems, or large-scale research workflows.
  • Experience with Git-based workflows and internal developer tooling.
  • Prior experience in environments where reliability, scale, and speed all matter.

About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Information Technology Jobs

Find similar Software Engineer - Data Aquisition (systems) jobs: