Full Job Description
Please note this posting is to advertise potential job opportunities. This exact role may not be open today but could open in the near future. When you apply, a Cisco representative may contact you directly if a relevant position opens.
As part of this role, you will be responsible for maintaining services in a FedRAMP compliant environment, therefore, must be a U.S. citizen. This position may also perform work that the U.S. government has specified can only be performed by a U.S. citizen on U.S. soil.
Applications are accepted until further notice.
Your Impact
Monitor and triage infrastructure and application alerts across our multi-tenant cloud environments to maintain efficient platform uptime and reliability.
• Lead end-to-end incident response and major incident management workflows to minimize Mean Time to Resolution (MTTR) during service-impacting events.
• Implement and refine automated playbooks and standard operating procedures to streamline operational handoffs and eliminate repetitive solving steps.
• Collaborate closely with Site Reliability Engineering (SRE) and Product teams during Post-Incident Reviews (PIRs) to identify root causes and implement permanent architectural remediations.
• Participate in an operational on-call rotation to guarantee round-the-clock stability and seamless customer experiences across global production clusters
Minimum Qualifications:
• Meet one of the following (completed within the past 3 years or expected to be completed within the next 12 months):
-A Master's degree with 0 years of work experience; or
-A Bachelor's degree with 2 years of work experience; or
-An educational or workforce development program (e.g., Associate's, Apprenticeship, Boot camp, or Certification) with 3 years of work experience; or
-A High school diploma (or equivalent) with 4 years of work experience.
Qualifying education and work experience should be inComputer Science, Information Technology, or a related field.
Preferred Qualifications
• Exposure with configuration management and infrastructure-as-code automation tools such as Puppet, Ansible, or Terraform.
• Foundational knowledge of containerization and orchestration technologies, including Docker and Kubernetes.
• Experience using Splunk (SPL, dashboards, alerting) or OpenTelemetry frameworks to analyze logs, traces, and metrics for operational anomalies.
• Experience utilizing AI developer tooling (such as Claude, GitHub Copilot, or Codex) to accelerate script development, log parsing, and operational runbook generation.