Job Description:The Hivemind Software Engineering Integration and Test team is seeking a Staff Automated Test Engineer to provide technical leadership for automated verification across our next-generation autonomy platform. You will define and evolve test architecture, validation infrastructure, MLOps quality systems, and CI/CD pipelines that enable the Hivemind software ecosystem to create, test, and deploy resilient autonomy capabilities for unmanned aircraft and robotic platforms operating in complex, contested, and GPS-denied environments.
You will work across production flight code, machine learning models, simulation and synthetic environments, mission planning and orchestration systems, operator-facing Ground Control Station applications, telemetry and data pipelines, cloud-native developer infrastructure, and hardware-in-the-loop systems.
In this hands-on Staff role, you will lead complex, cross-functional initiatives, establish automation and verification standards, identify systemic quality risks, and develop scalable test infrastructure. The ideal candidate combines deep Python automation and distributed-systems testing experience with technical leadership and experience validating machine learning systems throughout the data, training, evaluation, deployment, and monitoring lifecycle.
What you'll do:- Own the technical strategy, architecture, and roadmap for automated testing, verification, and MLOps quality across the Hivemind ecosystem.
- Design and maintain scalable test frameworks for autonomy software, backend services, APIs, operator-facing applications, and distributed hardware environments.
- Lead functional, integration, regression, system, performance, reliability, and end-to-end testing across simulation, edge-compute, software-in-the-loop, and hardware-in-the-loop environments.
- Build automated ML validation pipelines covering data quality, training reproducibility, model accuracy, robustness, regression, latency, resource utilization, and system integration.
- Establish CI/CD and continuous training workflows that provide versioning and traceability for datasets, models, configurations, evaluation results, and deployment artifacts.
- Develop scenario-based validation for autonomy models, including edge cases, degraded sensing or communications, distribution shifts, and representative mission conditions.
- Create observability, analytics, and failure-triage capabilities for software behavior, model and data drift, inference health, test results, and production performance.
- Build Python automation that improves test execution, parallelization, reporting, environment setup, experiment comparison, and developer productivity.
- Create test harnesses, simulators, stubs, mocks, and synthetic data capabilities that improve system testability and coverage.
- Collaborate with software, autonomy, machine learning, data, simulation, and systems engineers to define verification strategies and improve designs before implementation.
- Develop and govern AI-assisted engineering workflows using coding agents and LLM-based tools for test generation, log analysis, debugging, and failure triage while maintaining security, reproducibility, and traceability.
Required qualifications:- Typically 8+ years of relevant experience in software engineering, test infrastructure, developer tooling, MLOps, systems integration, or systems verification, or an equivalent combination of experience and demonstrated impact.
- 5+ years of experience building scalable automation frameworks or developer tooling in Python.
- Demonstrated success designing test, CI/CD, or MLOps infrastructure used across multiple engineering teams.
- Experience validating machine learning systems across the data, training, evaluation, packaging, deployment, and monitoring lifecycle.
- Understanding of ML quality risks such as data leakage, training-serving skew, nondeterminism, distribution shift, drift, model regression, and statistical acceptance criteria.
- Experience defining model-performance baselines, automated evaluation suites, release thresholds, and candidate-to-production comparison workflows.
- Experience testing GPU-accelerated infrastructure and workloads, including GPU scheduling, allocation, utilization, and resource contention in Kubernetes environments.
- Experience with performance benchmarking, profiling, and observability for GPU workloads, including identifying compute, memory, storage, networking, and data-loading bottlenecks.
- Experience validating multi-tenant Kubernetes environments, including RBAC, resource quotas, workload isolation, and scheduling behavior.
- Experience qualifying integrated hardware and software systems, including automated validation of compute, GPU, storage, networking, drivers, firmware, and deployed software configurations.
- Strong system-design skills and experience testing distributed systems, backend services, APIs, and integrated hardware and software environments.
- Experience developing integration and regression strategies for internally developed, third-party, open-source, and partner software, including dependency management, compatibility testing, and upgrades.
- Strong understanding of asynchronous and concurrent Python programming for scalable automation and parallel test execution.
- Experience with package and dependency management, reproducible environments, and build systems such as Conan, pip, setuptools, Poetry, Nix, or similar.
- Experience with automated observability, log collection, analytics, reporting, and root-cause analysis in complex software, data, and infrastructure systems.
- Experience working in Linux-based development environments.
Preferred qualifications:- Experience with model registries, experiment tracking, dataset or feature versioning, model serving, and automated artifact promotion using MLflow, Kubeflow, Weights & Biases, SageMaker, Vertex AI, or similar platforms.
- Experience with GPU scheduling and orchestration platforms such as Run:ai, NVIDIA GPU Operator, KAI Scheduler, Kueue, Volcano, or similar technologies.
- Experience with NVIDIA GPU infrastructure, including CUDA, drivers, container runtimes, Multi-Instance GPU, GPU fractionalization, and hardware/software compatibility testing.
- Experience with GPU profiling and performance-analysis tools such as NVIDIA Nsight, PyTorch Profiler, or similar technologies.
- Experience validating perception, planning, decision-making, reinforcement learning, or other autonomy models in simulation and on deployed systems.
- Experience testing models on embedded or edge-compute platforms, including latency, memory, power, accelerator compatibility, quantization, and hardware-specific behavior.
- Experience qualifying production servers or appliances, including hardware validation, burn-in, provisioning, firmware, networking, storage, and software-stack validation before deployment.
- Experience with reliability, fault-injection, and recovery testing across distributed compute, storage, networking, and GPU infrastructure.
- Experience validating reproducible installation, operation, upgrades, and rollback in cloud, on-premises, disconnected, or air-gapped environments.
- Experience with containers, Kubernetes, cloud infrastructure, infrastructure as code, and reproducible test environments.
- Proficiency with Go or TypeScript for automation tooling or UI test development.
- Experience integrating Python with native C or C++ applications through bindings, wrappers, subprocess interfaces, or similar interoperability tooling.
- Aerospace, robotics, autonomy, embedded systems, or safety-critical software experience.
- Familiarity with software-in-the-loop, hardware-in-the-loop, requirements-based verification, configuration management, artifact traceability, or standards such as DO-178C and MIL-STD-882.
$150,000 - $230,000 a year
#LI-DS5
#LD
Full-time regular employee offer package:
Pay within range listed + Bonus + Benefits + Equity
Temporary employee offer package:
Pay within range listed above + temporary benefits package (applicable after 60 days of employment)
Salary compensation is influenced by a wide array of factors including but not limited to skill set, level of experience, licenses and certifications, and specific work location. All offers are contingent on a cleared background and possible reference check. Military fellows and part-time employees are not eligible for benefits. Please speak to your talent acquisition representative for more information.