Position OverviewYou are the first human in the loop - the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations.
NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans - it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM-8PM PST shift. You execute SOPs, escalate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents.
What you'll own- Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
- Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
- Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
- Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
- Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
- Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
- Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
- Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
- Maintain and update operational runbooks based on recurring issues.
- Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
Feed the AIOps substrate- Every novel incident you resolve is data the platform team needs - you tag it, describe it, and hand it back so it becomes an automation.
- Every runbook you touch should get closer to being executable by the platform, not by you.
- Your handoff notes are structured signal, not free-form email.
Why this role is different from a NOC job- You are not the last line of defense - the platform is. You are the training signal.
- Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors.
Job Requirement:- 2+ years in NOC, data center operations, or IT support role
- Basic Linux system administration (command line, log analysis, service management)
- Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
- Experience with ticketing systems (ServiceNow, Jira Service Management)
- Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
- Strong communication skills for shift handoffs, incident documentation, and escalation
- Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
- Curiosity about automation - you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
- Comfort with structured data - you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.