Full Job Description
We are seeking an experienced Senior Software Engineer to contribute to the design, development, and upstream enablement of Cornelis Networks' AI/HPC communication middleware. This role focuses on enabling and optimizing HPC middleware (MPI/SHMEM) and AI middleware (collective communication libraries / CCL, e.g., NCCL/RCCL and related stacks) over Cornelis Networks' interconnects. You will help drive performance, correctness, and deployability across real customer workloads by collaborating across the stack (drivers/kernel, switches, systems, and application/framework teams) and by contributing upstream to key open-source projects.
Key Responsibilities
• AI/HPC Middleware Enablement & Optimization: Implement and help optimize features for HPC middleware (HPI and SHMEM) and AI middleware (CCL stacks ie: NCCL/RCCL and related collective communication libraries)
• Transport & Integration: Assist in integrating middleware capabilities with underlying transports and provider layers (ie: libfabric / OFI, UCX, verbs-like semantics where applicable)
• Upstream & Open-Source Leadership: Prepare and submit upstream contributions across MPI/SHMEM projects, CCL ecosystems, and related components
• Cross-Stack Collaboration & Performance Validation: Collaborate with kernel/driver and switch teams to help deliver end-to-end performance aligned to the Cornelis product roadmap
• Customer & Field Impact: Help analyze performance traces and reproduce customer issues, and assist in translating findings into fixes.
Minimum Qualifications
• 6+ years of experience in systems programming in C/C++ on Linux; Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience). Master's or Doctorate a plus.
• Hands-on experience with HPC middleware internals (Open MPI, MPICH, MVAPICH, and/or SHMEM) or optimizing collective communications for AI, including NCCL/RCCL, CUDA/ROCm) and/or other CCL stacks.
• Demonstrated ability to deliver low-latency/high-throughput communication paths and to diagnose performance issues using profiling/tracing tools.
• Working knowledge of transport/integration layers such as OFI/libfabric, UCX, verbs-style concepts.
• Solid background in RDMA concepts and performance tuning.
• Experience contributing to open-source projects.
Preferred Qualifications
• Prior experience developing or maintaining libfabric providers.
• Familiarity with Ultra Ethernet (UEC/UET) specifications.
• Experience with RoCEv2, congestion control, and/or Ethernet-based RDMA deployments.
• Experience with cluster-scale benchmarking, profiling, and optimization.
• Background with HPC fabrics, such as Omni-Path.
Location: This is a remote position for employees residing within the United States.
We offer a competitive compensation package that includes equity, cash, and incentives, along with health and retirement benefits. Our dynamic, flexible work environment provides the opportunity to collaborate with some of the most influential names in the semiconductor industry.
At Cornelis Networks your base salary is only one component of your comprehensive total rewards package. Your base pay will be determined by factors such as your skills, qualifications, experience, and location relative to the hiring range for the position. Depending on your role, you may also be eligible for performance-based incentives, including an annual bonus or sales incentives.
In addition to your base pay, you'll have access to a broad range of benefits, including medical, dental, and vision coverage, as well as disability and life insurance, a dependent care flexible spending account, accidental injury insurance, and pet insurance. We also offer generous paid holidays, 401(k) with company match, and Open Time Off (OTO) for regular full-time exempt employees. Other paid time off benefits include sick time, bonding leave, and pregnancy disability leave.