Senior AI Infrastructure Engineer

Tencent

• $124K — $283K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • 5+ years of experience in data center infrastructure, deployment, or operations.
  • Solid understanding of high-performance compute server hardware and networking fundamentals.
  • Proven experience evaluating data center facilities and reviewing technical proposals.
  • Strong track record coordinating complex technical projects with external partners.
  • Hands-on experience with cluster orchestration tools and Linux system administration.
  • Proficiency in scripting/automation with Python or Bash.

Responsibilities

  • Conduct on-site data center assessments for AI infrastructure requirements.
  • Review and validate AI infrastructure solution designs, identifying risks and trade-offs.
  • Coordinate with external partners to drive timelines and resolve technical issues.
  • Translate internal requirements into technical specifications with business and engineering teams.
  • Participate in day-2 operations for live AI environments, including monitoring and incident response.
  • Build and improve operational standards and documentation for infrastructure delivery.
  • Provide technical guidance and mentorship to junior engineers and interns.

Benefits

  • Medical, dental, vision, life, and disability benefits.
  • Participation in the company's 401(k) plan.
  • Up to 15 to 25 days of vacation per year based on tenure.
  • Up to 13 holidays throughout the year.
  • Up to 10 days of paid sick leave per year.
Full Job Description
What the Role Entails
Role Summary

We are looking for an experienced Senior AI Infrastructure Engineer to join our team. This role owns the technical evaluation and end-to-end execution of AI infrastructure deployments - from data center due diligence and solution review, through cross- functional delivery coordination, to ongoing operations of production AI environments. You will act as a key technical owner bridging internal stakeholders and external partners, ensuring infrastructure is delivered on time, to spec, and operated reliably at scale.Key Responsibilities
• Conduct on-site data center assessments to evaluate whether candidate facilities meet AI infrastructure requirements.
• Review and validate AI infrastructure solution designs, identifying technical risks, gaps, and cost/performance trade-offs before sign-off.
• Coordinate with external partners, driving timelines, resolving technical issues, and ensuring deliverables meet internal requirements.
• Partner with internal business and engineering teams to translate requirements into deliverable technical specifications.
• Participate in and eventually take ownership of day-2 operations for live AI

environments - monitoring, incident response, capacity/health checks, firmware and lifecycle management, coordinating hardware maintenance as needed.
• Build and improve operational standards, runbooks, and documentation for

infrastructure delivery and operations to enable consistent execution across regions.
• Provide technical guidance and mentorship to junior engineers/interns on hardware diagnostics, cluster tooling, and best practices.
• Track and report on project status and risks to management.

Who We Look For
Requirements
• Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
• 5+ years of experience in data center infrastructure, infrastructure deployment, or infrastructure operations.
• Solid understanding of high-performance compute server hardware, high-speed

networking (e.g., InfiniBand/RoCE), storage systems, and data center power & cooling fundamentals.
• Proven experience evaluating data center facilities and reviewing technical proposals for compute infrastructure.
• Strong track record coordinating complex, multi-party technical projects to closure, including working effectively with external partners and cross-regional teams.
• Hands-on experience with cluster orchestration and management tools (e.g., Slurm, Kubernetes) and Linux system administration.
• Proficiency in scripting/automation (Python, Bash; Ansible/Terraform a plus).
• Familiarity with observability/monitoring stacks (Prometheus, Grafana, ELK) and DCIM tooling.
• Bilingual proficiency in English and Mandarin highly preferred
• Strong ownership mentality, able to independently drive projects with minimal

supervision.
Preferred Qualifications
• Experience with high-performance storage / parallel file systems (e.g., Lustre,

GPFS/Spectrum Scale, WekaFS, VAST, Ceph) in AI infrastructure environments.
• Experience deploying/operating large-scale AI compute environments in a hyperscale or cloud environment.
• Data center or infrastructure certifications.

Location State(s)

US-California-Palo Alto

The expected base pay range for this position in the location(s) listed above is $124,800.00 to $283,800.00 per year. Actual pay may vary depending on job-related knowledge, skills, and experience.Employees hired for this position may be eligible for a sign on payment, relocation package, and restricted stock units, which will be evaluated on a case-by-case basis.Subject to the terms and conditions of the plans in effect, hired applicants are also eligible for medical, dental, vision, life and disability benefits, and participation in the Company's 401(k) plan. The Employee is also eligible for up to 15 to 25 days of vacation per year (depending on the employee's tenure), up to 13 days of holidays throughout the calendar year, and up to 10 days of paid sick leave per year.Your benefits may be adjusted to reflect your location, employment status, duration of employment with the company, and position level. Benefits may also be pro-rated for those who start working during the calendar year.

Similar Jobs

More Jobs at Tencent

More Information Technology Jobs

Find similar Senior AI Infrastructure Engineer jobs: