Cisco

Senior Site Reliability Engineer (SRE) (Hybrid)

Cisco • $167K — $245K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years' experience in Site Reliability Engineering, Platform, Cloud, Infrastructure Engineering, or related fields
  • 3+ years' experience operating Kubernetes in production environments
  • Proficiency in building and maintaining CI/CD platforms and deployment automation
  • Experience with major cloud platforms such as AWS or GCP

Responsibilities

  • Operate and enhance Kubernetes-based production infrastructure
  • Manage customer deployments in both cloud and air-gapped environments
  • Develop and improve deployment observability, monitoring, logging, and alerting systems
  • Automate to boost operational efficiency and performance optimization
  • Participate in incident response and root cause analysis to improve reliability
  • Tune infrastructure components to enhance performance and resilience
  • Design internal tooling using Python and/or Go

Benefits

  • Hybrid work model with 2 days on-site per week
  • Opportunity to work with cutting-edge AI resilience technology
  • Collaboration with engineering teams to enhance reliability and security
  • Emphasis on professional development and operational excellence
Full Job Description
The application window is expected to close on: 09/29/2026
This is a Hybrid position requiring approximately 2 days per week on-site at Cisco offices in either San Francisco, San Jose, or New York City.

The Splunk Agent Resilience team is defining the future of AI resilience. Our team provides scalable, cost-effective evaluation and guardrails that ensure AI agents behave as intended, improving reliability and reducing risks. This unified approach empowers our customers to confidently deploy and manage AI-powered applications with enhanced observability and control.

As a Senior Site Reliability Engineer (SRE), you will build, operate, and continuously improve the reliability, scalability, and operational excellence of Splunk Agent Observability's deployment platform and production infrastructure. You will own the operational backbone supporting both cloud and air-gapped customer deployments, develop automation to improve operational efficiency, and partner closely with engineering teams to deliver reliable, secure, and scalable systems.

Your Impact
  • Operate and improve Kubernetes-based production infrastructure and deployment systems
  • Own customer deployments across cloud and air-gapped environments, including installation, upgrades, troubleshooting, and lifecycle management
  • Build and improve deployment observability, monitoring, logging, and alerting
  • Improve reliability, scalability, and operational efficiency through automation and performance optimization
  • Participate in production incident response, root cause analysis, and reliability improvements
  • Tune infrastructure components-including databases and services-to improve performance and resilience
  • Design and develop internal tooling using Python and/or Go
  • Debug complex production issues spanning Kubernetes, networking, storage, and application layers
  • Manage infrastructure using Terraform or similar Infrastructure as Code tools
  • Collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures

Minimum Qualifications
  • 7+ years' experience with a Bachelor's degree or 4+ yrs with a Masters or 1 year with a PhD, or equivalent related experience; Experience to include at least 4 years in Site Reliability Engineering, Platform / Cloud / Infrastructure Engineering, or related fields
  • 3+ years' operating Kubernetes in production; experience with Helm
  • Experience building and maintaining CI/CD platforms and deployment automation
  • Experience working with AWS, GCP, or similar cloud platforms

Preferred Qualifications
  • Experience improving production reliability, scalability, and availability
  • Experience with monitoring, logging, observability, and alerting platforms
  • Strong scripting or programming skills in Python and/or Go
  • Experience with Infrastructure as Code tools such as Terraform
  • Experience with MLOps preferred
  • Solid understanding of networking fundamentals (VPCs, DNS, routing, load balancing)
  • Experience operating databases and tuning them for performance and reliability
  • Experience deploying and supporting cloud and air-gapped/on-prem environments
  • Strong debugging and troubleshooting skills across distributed systems starts with you.

About Cisco

Cisco Careers

Join the vibrant team at Cisco, a global leader in networking and cybersecurity solutions, where innovation and leadership thrive. Cisco offers a plethora of job opportunities that cater to a range of skills and experiences, making it an ideal place for both seasoned professionals and those seeking an internship to jumpstart their career. Work You’ll Do At Cisco, you’ll be part of a culture that values diversity, leadership, and professional growth. Engage in work that matters with a team that combines technology, creativity, and the power of human connection to redefine networking. Cisco’s commitment to innovation isn’t just about technology, but also about transforming the way we work and collaborate. Cisco’s employment philosophy supports career advancement and nurtures a leadership pipeline that is equipped with diversity training and opportunities for growth. Whether you’re applying your skills to drive our latest innovations or using our vast networking capabilities to solve complex problems, at Cisco, every role is impactful. Join Our Dynamic Team Explore job opportunities in areas ranging from engineering to marketing, sales to cybersecurity. Cisco is hiring individuals who are passionate, curious, and ready to drive change. Positions at Cisco offer competitive benefits, a supportive culture, and the chance to work with cutting-edge technology. Internship Programs Kickstart your career with a Cisco internship. Gain invaluable industry experience, enhance your resume, and build professional networks that last a lifetime. Our internships provide hands-on experience and the chance to work on projects that matter. Leadership and Development Cisco is committed to fostering leadership skills and providing employees with the training needed to succeed. Our leadership programs help you develop new skills, manage teams effectively, and lead with confidence. Cisco’s commitment to professional development ensures that your career path is as dynamic as our technologies. Benefits and Culture Cisco understands the importance of a balanced life. Our benefits package is designed to ensure that our team members are healthy, happy, and secure. At Cisco, you’ll find a supportive culture that encourages open communication, teamwork, and mutual respect. Stay Connected Join Cisco’s Talent Network Stay informed about new positions that match your skills and interests. At Cisco, we value the curiosity and unique perspectives of our team members. Subscribe to receive personalized job alerts and insider tips directly from our hiring managers. Explore Cisco Jobs Ready to advance your career at Cisco? Search open positions, prepare your resume, and get ready for an interview that could lead to your next big opportunity. At Cisco, we’re not just filling positions—we’re investing in leaders. Keep Up to Date Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here. READ CAREERS BLOG Job Alert Emails Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await you at Cisco.
Learn more about Cisco
Size
79,500 employees
Market Cap
$194.5 billion
Industry
Net Income
$10.1 billion
Founded
2014
5 Year Trend
+1.4%
Revenue
$48 billion
NASDAQ

Similar Jobs

More Jobs at Cisco

More Information Technology Jobs

Find similar Senior Site Reliability Engineer (SRE) (Hybrid) jobs: