Infrastructure Software Engineer, Fleet & Automation

Nscale

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent experience.
  • 5+ years of experience in building large-scale infrastructure applications.
  • Proficient in programming languages including C, C++, Java, and Python for API design and testing.
  • Strong understanding of Linux OS, networking (TCP/IP, BGP), and configuration management tools like Ansible and Terraform.
  • Experience with distributed systems and infrastructure tooling such as DCIMs, OpenStack, and bare metal APIs.

Responsibilities

  • Develop and implement the technical architecture for workflow automation systems.
  • Oversee comprehensive workflows for device provisioning and testing at scale.
  • Design systems for hardware lifecycle management, including GPU nodes.
  • Collaborate with Infrastructure and SRE teams to build scalable automation solutions.
  • Set engineering standards focused on reliability and operational excellence.
  • Create production-grade Python automation systems for hardware lifecycle tasks.
  • Work with cross-functional teams to develop maintainable automated systems.

Benefits

  • Competitive compensation package with equity options and annual reviews.
  • Opportunity to work in a rapidly growing tech startup focused on AI.
  • Personalized progression plan that supports career growth and innovation.
Full Job Description
Overview

As an Infrastructure Software Engineer for Fleet & Automation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering principles, you will focus on building and maintaining the control plane, tooling, and automation that supports Fleet Operations, Network Operations, and Observability functions. Your work will directly translate into higher system availability and reduced operational costs.
Key Responsibilities
  • Perform technical architecture, roadmap and implementation for workflow automation systems, driving architecture decisions that balance automation complexity, reliability, and maintainability. Identify and resolve performance and scalability issues. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
  • Own end-to-end delivery of device provisioning, validation, testing, and remediation workflows at scale.
  • Design and build workflow orchestration systems for hardware lifecycle management, including GPU nodes and network switches.
  • Partner with Infrastructure, Platform, and SRE teams to translate operational needs into robust, scalable automation.
  • Establish engineering standards for reliability, observability, and operational excellence across all services. Help set up engineering best practices in collaboration with the broader engineering team.
  • Build production-grade Python systems for hardware lifecycle automation, leveraging AI tools to accelerate delivery. Assess impact to team software stack from new hardware product programs and explore AI driven process improvement and automation.
  • Collaborate with cross-functional teams (product, design, operations, infrastructure) to build efficient, interoperable, and maintainable automated systems.
Required Qualifications
  • Education: Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • Experience: 5+ years relevant experience building large-scale infrastructure applications or similar experience.
  • Programming: Experience in utilizing languages such as C, C++, Java, and scripting languages such as Python for API design and unit testing techniques.
  • Systems Expertise: Deep understanding of Linux operating systems, networking fundamentals (TCP/IP, BGP), and familiarity with configuration management tools (e.g., Ansible, Terraform).
  • Distributed Systems: Experience building, running and debugging large-scale infrastructure, stateful and stateless services for distributed systems or networks, and experience with compute technologies, storage, or hardware architecture. Experience integrating with infrastructure tooling such as: DCIMs, NetBox, OpenStack, bare metal APIs (MAAS, Ironic, IPMI).
Preferred Qualifications
  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • Experience designing, analyzing and improving efficiency, scalability, and performance of various system resources.
  • Direct experience with AI/HPC infrastructure, including NVIDIA GPUs, InfiniBand or high-speed Ethernet fabrics, and related management software (e.g., NCCL, SLURM).
  • Experience with advanced observability and monitoring systems (Prometheus, Grafana, OpenTelemetry) for complex, high-cardinality telemetry data.
  • Familiarity with cloud-native technologies (Kubernetes, Docker) and infrastructure-as-code principles.
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
  • Familiarity with SLOs/metrics measurement, logs/telemetry/metrics integration with tools for enhanced operator experience.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.

Similar Jobs

More Jobs at Nscale

More Information Technology Jobs

Find similar Infrastructure Software Engineer, Fleet & Automation jobs: