Infrastructure Software Engineer, Fleet & Automation

Nscale

$135K — $160K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or similar field, or equivalent experience.
  • Over 5 years of experience in building large-scale infrastructure applications or similar.
  • Proficient in programming languages such as C, C++, Java, and Python for API design and unit testing.
  • Strong knowledge of Linux OS, network fundamentals (TCP/IP, BGP), and config management tools like Ansible and Terraform.
  • Experienced with distributed systems, including stateful and stateless services, compute technologies, and infrastructure tooling.

Responsibilities

  • Develop architecture and roadmap for workflow automation, balancing complexity and reliability.
  • Drive the end-to-end process for device provisioning and testing at scale.
  • Create workflow orchestration systems for managing hardware lifecycles, including GPUs and network switches.
  • Collaborate with teams to turn operational needs into efficient automation solutions.
  • Set engineering standards for reliability and observability across services, establishing best practices.
  • Build production-grade Python systems for hardware automation using AI tools.
  • Engage with cross-functional teams to develop maintainable automated systems.

Benefits

  • Collaborative and innovative work environment with real impact.
  • Competitive compensation package with annual reviews.
  • Opportunity to work at a rapidly growing tech startup and engage with leading minds.
  • Personalized progression plan to support career growth and exploration.
Full Job Description
Overview

As an Infrastructure Software Engineer for Fleet & Automation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering principles, you will focus on building and maintaining the control plane, tooling, and automation that supports Fleet Operations, Network Operations, and Observability functions. Your work will directly translate into higher system availability and reduced operational costs.
Key Responsibilities
  • Perform technical architecture, roadmap and implementation for workflow automation systems, driving architecture decisions that balance automation complexity, reliability, and maintainability. Identify and resolve performance and scalability issues. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
  • Own end-to-end delivery of device provisioning, validation, testing, and remediation workflows at scale.
  • Design and build workflow orchestration systems for hardware lifecycle management, including GPU nodes and network switches.
  • Partner with Infrastructure, Platform, and SRE teams to translate operational needs into robust, scalable automation.
  • Establish engineering standards for reliability, observability, and operational excellence across all services. Help set up engineering best practices in collaboration with the broader engineering team.
  • Build production-grade Python systems for hardware lifecycle automation, leveraging AI tools to accelerate delivery. Assess impact to team software stack from new hardware product programs and explore AI driven process improvement and automation.
  • Collaborate with cross-functional teams (product, design, operations, infrastructure) to build efficient, interoperable, and maintainable automated systems.
Required Qualifications
  • Education: Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • Experience: 5+ years relevant experience building large-scale infrastructure applications or similar experience.
  • Programming: Experience in utilizing languages such as C, C++, Java, and scripting languages such as Python for API design and unit testing techniques.
  • Systems Expertise: Deep understanding of Linux operating systems, networking fundamentals (TCP/IP, BGP), and familiarity with configuration management tools (e.g., Ansible, Terraform).
  • Distributed Systems: Experience building, running and debugging large-scale infrastructure, stateful and stateless services for distributed systems or networks, and experience with compute technologies, storage, or hardware architecture. Experience integrating with infrastructure tooling such as: DCIMs, NetBox, OpenStack, bare metal APIs (MAAS, Ironic, IPMI).
Preferred Qualifications
  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • Experience designing, analyzing and improving efficiency, scalability, and performance of various system resources.
  • Direct experience with AI/HPC infrastructure, including NVIDIA GPUs, InfiniBand or high-speed Ethernet fabrics, and related management software (e.g., NCCL, SLURM).
  • Experience with advanced observability and monitoring systems (Prometheus, Grafana, OpenTelemetry) for complex, high-cardinality telemetry data.
  • Familiarity with cloud-native technologies (Kubernetes, Docker) and infrastructure-as-code principles.
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
  • Familiarity with SLOs/metrics measurement, logs/telemetry/metrics integration with tools for enhanced operator experience.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.

Similar Jobs

More Jobs at Nscale

More Information Technology Jobs

Find similar Infrastructure Software Engineer, Fleet & Automation jobs: