Software Engineer - Data Center

SpaceXAI

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field.
  • 3+ years of experience in building and operating production software.
  • Strong knowledge of data structures, algorithms, operating systems, and networking.
  • Proficiency in a programming language like Rust, Python, Java, JavaScript, or C++.
  • Experience designing and developing RESTful APIs for critical applications.
  • Familiarity with databases such as Postgres, MongoDB, MySQL, or DynamoDB.
  • Experience working in cloud environments like GCP, Azure, AWS, or OCI.

Responsibilities

  • Build and operate scalable and auditable software stacks for site operations.
  • Design, operate, and maintain multi-service production systems.
  • Ensure accuracy of operational state and data integrity.
  • Develop integrations with ticketing, inventory, and telemetry systems.
  • Maintain tool reliability through access control and safe deployments.
  • Engage with SiteOps and NOC users to measure workflow adoption.

Benefits

  • Fully onsite role in Memphis, TN or Southhaven, MS.
  • Opportunity to work on high-leverage tools for data center operations.
  • Collaborative environment with cross-functional teams.
Full Job Description
ABOUT THE ROLE:

The Data Center Engineering team builds the internal systems and platforms that keep our data centers running at the scale and reliability required for frontier AI training and inference. We partner closely with datacenter operations, research, and infrastructure teams to deliver high-leverage tools that turn raw operational data into clear insight and action.
RESPONSIBILITIES:
  • Build and operate the software stacks that make site operations scalable, auditable, and fast - including repair trackers, vendor turnback workflows, operational dashboards, and their integrations. Your users are the technicians, managers, NOC operators, and leadership who run the fleet.
  • Design, build, and operate multi-service production systems (UI, APIs, data pipelines, auth, and on-call) for systems such as SRT-style repair/maintenance trackers and vendor turnover/turnback state machines.
  • Own correctness of operational state: node state accuracy, queue ownership, and audit trails - the data that decides what work happens on the floor.
  • Build and maintain integrations with ticketing, inventory/rack systems, telemetry stores, and vendor portals.
  • Keep the tools themselves reliable: uptime, data integrity, reconciliation, access control, and safe deploys.
  • Embed with SiteOps and NOC users; measure workflow adoption, not just feature delivery.
BASIC QUALIFICATIONS:
  • Bachelor's degree in Computer Science, Engineering, or related fields.
  • 3+ years building and operating production software (backend and/or full-stack).
  • Strong fundamental knowledge of computer science - data structures, algorithms, operating systems and networking.
  • Strong proficiency in at least one programming language e.g. Rust, Python, JavaScript, Java, C++, etc.
  • Experience designing and developing RESTful APIs for mission-critical applications.
  • Experience working with at least one database like Postgres, MongoDB, MySQL, DynamoDB, etc.
  • Experience collaborating with cross-functional teams.
  • Experience working in cloud platforms like GCP, Microsoft Azure, AWS, OCI or similar.
  • Experience using observability tools and dashboards like New Relic, Splunk, Grafana, etc.
  • Experience writing unit tests and integration tests.
PREFERRED SKILLS AND EXPERIENCE:
  • Full-stack or backend + data experience, especially with workflow and state-machine systems.
  • Experience automating deployments using CI/CD tools like Azure Devops, ArgoCD, Jenkins, GitHub Actions, etc.
  • Experience building high-correctness operational UIs where a wrong value can dispatch a human to the wrong rack.
  • On-call discipline and a track record of treating internal platforms with production rigor.
  • Experience designing and shipping multi-service systems (APIs, data stores, and at least one of: UI, pipelines, or auth).
  • Experience delivering high-quality internal tools in rapidly changing environments (as a tech lead, former founder, etc.).
  • Proven ownership of correctness-sensitive systems - state machines, workflows, or operational data where inaccurate state has real-world impact.
  • Experience integrating with external systems via APIs (e.g. ticketing, inventory, telemetry, or vendor portals).
  • MS in Computer Science or related field.
  • Experience collaborating closely with operations, NOC, and infrastructure teams.
  • Experience working with performance load testing tools like BlazeMeter, k6, etc.
ADDITIONAL REQUIREMENTS:
  • The role is fully onsite in Memphis, TN or Southhaven, MS. Candidates are expected to be located near the area or open to relocation.


Similar Jobs

More Jobs at SpaceXAI

More Information Technology Jobs

Find similar Software Engineer - Data Center jobs: