Super Micro Computer, Inc

Sr. System Engineer

Super Micro Computer, Inc$137K — $156K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS/MS in Electrical or Computer Engineering, with MS preferred
  • 8+ years in server/network/storage hardware configuration and troubleshooting
  • 8+ years in DevOps or cloud environments (Docker/Containers, Kubernetes)
  • Experience with leading AI/ML frameworks (PyTorch, TensorFlow)
  • Proficient in TCP/IP protocols and application protocols
  • Familiar with HPC, AI or Cloud benchmark tests
  • Strong programming skills in Python and shell scripting

Responsibilities

  • Deploy rack/cluster infrastructure and conduct system-level testing using proprietary tools
  • Conduct proof of concept design and testing, optimizing HPC/AI applications
  • Lead daily operational support for cloud and cluster infrastructure while documenting issues
  • Write technical documents for testing and troubleshooting related to hardware and software
  • Deliver on-site deployment services, ensuring customer satisfaction
  • Develop automation tools for cluster deployment and testing environments

Benefits

  • Comprehensive benefits package
  • Eligibility for bonus and equity award programs
  • Opportunities for professional development
  • Collaborative and innovative work environment
  • Focus on cutting-edge technologies in HPC and AI solutions
Full Job Description
Job Req ID: 29799

Job Summary:

As a global leader in server technologies, Supermicro has been growing extremely fast in many key markets such as Cloud Computing, Big Data, HPC, AI and Storage, etc. To meet the market demand, Supermicro is developing end to end enterprise IT solutions with compute, storage, networking all integrated into full rack or multi-rack level systems. Senior System Engineer plays an important role in designing, implementing, testing and deploying rack system solutions for data center and enterprise customers.

Essential Duties and Responsibilities:

Includes the following essential duties and responsibilities (other duties may also be assigned):
• Deploy Rack/Cluster infrastructure and execute comprehensive system level testing on the latest GPUs, CPU processors, Network and Storage, encompassing functionality, compatibility, performance, stress, and reliability testing, leveraging proprietary in-house tools
• Conduct proof of concept design and testing. Establish expertise in HPC/AI applications and benchmarks, providing optimized benchmarks for HPC/AI applications by fine-tuning system settings, optimizing OS/network configurations, and demonstrating strong problem-solving skills and building robust processes and procedures for HPC/AI solutions
• Lead day-to-day operational support for Cluster, Storage, HPC and Cloud infrastructure. Identify and document hardware and software quality issues. Collaborate with product management and other Engineering teams to integrate enhancements into future products
• Write technical documents for test procedures, test reports and troubleshooting procedures related to servers/networks/clusters software and hardware to facilitate knowledge sharing
• Deliver on-site deployment services to ensure customer acceptance verification and satisfaction
• Write automation tools for cluster deployment and test environment

Qualifications:
• BS/MS in Electrical Engineering, Computer Engineering or a related field, MS preferred
• 8+ years of work-related experience in server/network/storage hardware configuration, testing, debugging and troubleshooting
• 8+ years of work-related experience in DevOps or in cloud environments, including but not limited to Docker/Containers and Kubernetes
• Experience with leading AI/ML frameworks such as PyTorch, TensorFlow, etc.
• Familiar with TCP/IP protocol stack, UDP, IPv4-IPv6, DNS, DHCP and other Application protocols
• Familiar with HPC, AI or Cloud benchmark tests, networking architecture
• Excellent Programming skills in Python and shell scripting
• Strong communication skills and strong sense of teamwork and good team player
• Familiar with MLPerf Training/Inference benchmark, LLM, HPL-AI or RCCL/NCCL is a plus
• CCNA, OpenStack, Openshit, Azure or AWS is a plus

Salary Range

$137,000 - $156,000

The salary offered will depend on several factors, including your location, level, education, training, specific skills, years of experience, and comparison to other employees already in this role. In addition to a comprehensive benefits package, candidates may be eligible for other forms of compensation, such as participation in bonus and equity award programs.

About Super Micro Computer, Inc

Super Micro Computer, Inc. designs, develops, manufactures, and sells server solutions based on modular and open architecture. The company offers a range of server, storage, and networking solutions, as well as application-optimized software. Super Micro Computer's products are used in data centers, cloud computing, enterprise IT, big data, high performance computing, and embedded systems. The company was founded in 1993 and is headquartered in San Jose, California.
Learn more about Super Micro Computer, Inc
Size
4,155 employees
Market Cap
$4.3 billion
Industry
Net Income
$88.5 million
5 Year Trend
+15.9%
Revenue
$3.2 billion
NASDAQ

Similar Jobs

More Jobs at Super Micro Computer, Inc

  • Super Micro Computer, Inc
    Sr. BIOS Engineer
    $140K — $170K *
    San Jose, CA 95123 (Santa Clara County)
    Enterprise Technology
    In-Person
  • Super Micro Computer, Inc
    Sr. System Engineer
    $137K — $156K *
    San Jose, CA 95123 (Santa Clara County)
    Information Technology
    In-Person
  • Super Micro Computer, Inc
    Data Analyst
    $90K — $110K *
    San Jose, CA 95123 (Santa Clara County)
    Information Technology
    In-Person
  • Super Micro Computer, Inc
    Manufacturing Engineer
    $73K — $100K *
    San Jose, CA 95123 (Santa Clara County)
    Manufacturing & Automotive
    In-Person
  • Super Micro Computer, Inc
    Program Manager - Infrastructure/Tools
    $100K — $125K *
    San Jose, CA 95123 (Santa Clara County)
    Technical Services
    In-Person

More Information Technology Jobs

Find similar Sr. System Engineer jobs: