Job Details:Job Description:We focus on designing, implementing, and optimizing AI accelerator systems and cloud infrastructure for large-scale machine learning workloads. The ideal candidate will have extensive experience with AI hardware platforms, system-level debugging, and cross-functional collaboration in enterprise environments.
Key Responsibilities:
AI/ML System Engineering
- Design and optimize AI accelerator systems (Gaudi, GPU clusters) for production ML workloads.
- Debug complex PCIe, memory subsystem, and interconnect issues in AI clusters.
- Validate and integrate the cutting-edge GPUs and AI accelerator platforms.
System Integration and Validation
- Lead platform bring-up and validation for next-generation AI hardware
- Develop comprehensive test plans for AI systems.
- Collaborate with OEM vendors on BMC firmware integration and system stability.
- Perform full-stack debugging across hardware, firmware, and software layers.
Infrastructure and Tooling
- Develop automated testing frameworks and monitoring solutions.
- Create diagnostic tools and APIs for system health monitoring.
Leadership and Collaboration
- Mentor junior engineers and data center technicians.
- Lead cross-functional teams through complex technical challenges.
- Coordinate with hardware, firmware, and software teams on platform readiness.
- Drive technical decisions and architectural improvements.
Qualifications:You must possess the below minimum qualifications to be initially considered for this position. Preferred qualifications are in addition to the minimum requirements and are considered a plus factor in identifying top candidates. Experience listed below would be obtained through a combination of your degree, research and or relevant previous job and or internship experiences.
Minimum Qualifications:- Bachelors and 6+ years or Masters and 4+ years or PhD and 2+ years in Computer Science, Electrical Engineering, or related field
- 5+ years of experience in system engineering, platform validation, or related roles.
- 3+ years of experience of successfully bringing up and debugging high-performance AI clusters.
- 3+ years of experience resolving complex system-level issues in production AI/ML environments.
- 3+ years of experience AI cluster design, validation, and production deployment experience.
Preferred Qualifications:- Experience with Intel platforms (Xeon, Gaudi) or similar GPU or AI accelerators.
- Familiarity with cloud deployment and containerization.
- Programming: Expert-level Python.
- AI/ML Frameworks: Experience with vLLM, PyTorch, TensorFlow, OpenMPI
- System Tools: Linux/Unix administration, Docker, shell scripting.
- Hardware: Deep understanding of PCIe, memory subsystems, AI accelerators.
- Protocols: Redfish, IPMI, BMC management.
- Computer architecture and microprocessor design.
- AI/ML workload optimization and deployment.
- System-level debugging and validation methodologies.
- Enterprise platform security and manageability.
Job Type:Experienced Hire
Shift:Shift 1 (United States of America)
Primary Location:US, Oregon, Hillsboro
Additional Locations:Position of TrustN/A
BenefitsWe offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel.
Annual Salary Range for jobs which could be performed in the US: $170,500.00-315,490.00 USD
The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.
Work Model for this RoleThis role will require an on-site presence. * Job posting details (such as work model, location or time type) are subject to change.
ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.