1. Collaborate with Cloud service team, internal teams, and cloud vendors to support infrastructure delivery and deployment activities.
2. Perform operating system reinstallation and reboot for cloud servers and bare metal environments.
3. Troubleshoot infrastructure issues involving Linux systems, cloud servers, networking, and hardware-related abnormalities.
4. Maintain operational data across internal systems including fault ticketing systems, repair records, and infrastructure tracking tools.
5. Work with cloud providers to resolve infrastructure incidents and large-scale operational issues.
6. Participate in on-call rotations and support incident management activities.
7. Monitor cloud infrastructure health, operational alerts, and asset utilization.
8. Develop or enhance operational tools, scripts, and automation workflows using Shell or Python.
9. Support operational process optimization and standardization initiatives.
10. Perform server and cloud network troubleshooting, including TCP/IP, VLAN, DNS, and Ipv6 related issues.
11. Support cloud fleet operations and infrastructure performance management.
12. Assist with documentation creation, knowledge sharing, and operational reporting.
Requirements
1. Bachelor’s Degree in Computer Science, Electrical Engineering, or related fields.
2. Experience in cloud infrastructure operations, server operations, or data center related environments.
3. Strong troubleshooting and analytical skills in Linux and infrastructure environments.
4. Experience with automated provisioning and OS deployment technologies such as PXE, iPXE, or provisioning pipelines.
5. Familiarity with public cloud platforms such as Oracle, Amazon Web Services, Google, or Microsoft.
6. Ability to understand, execute, and write Shell/Bash or Python scripts.
7. Familiarity with automation and infrastructure management tools such as Ansible, GitLab CI/CD, Terraform, or cloud SDKs.
8. Knowledge of Linux systems, cloud networking, and server lifecycle management.
9. Strong understanding of TCP/IP networking concepts including subnetting, VLANs, DNS, IPv6, and basic routing.
10. Familiarity with infrastructure monitoring, operational tooling, and incident management processes.
11. Strong communication, collaboration, and documentation skills.
12. Familiarity with large-scale cloud fleet operations is preferred.
13. Experience supporting GPU infrastructure, firmware lifecycle management, or RDMA networking is a plus