Cisco

Staff Site Reliability Engineer (SRE) (Hybrid)

Cisco$186K — $267K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in related fields with a Bachelor's degree; or 6+ years with a Master's; or 3+ years with a PhD.
  • 5+ years managing large-scale Kubernetes environments in production.
  • Expertise in designing resilient, distributed systems.
  • Familiarity with public cloud platforms (AWS, GCP, etc.).
  • Strong background in CI/CD and deployment automation.

Responsibilities

  • Define and implement the technical roadmap for platform reliability and scalability.
  • Oversee the evolution of deployment platforms for cloud and air-gapped environments.
  • Establish reliability standards like service level objectives (SLOs) and resiliency reviews.
  • Lead initiatives for improving reliability across Kubernetes and associated infrastructure.
  • Drive automation efforts to enhance operational efficiency and reduce manual tasks.
  • Design and develop internal tools for scalable, reliable operations.
  • Manage complex production incident responses and propose long-term solutions for recurring issues.
  • Collaborate with teams to create secure, scalable deployment architectures.

Benefits

  • Hybrid work model requiring 2 days on-site in major cities.
  • Opportunity to work with cutting-edge AI technologies.
  • Influence the technical roadmap and architecture of key platforms.
  • Engaging in major infrastructure initiatives with cross-team collaboration.
  • Mentoring and leadership opportunities across engineering teams.
Full Job Description
The application window is expected to close on: 08/15/2026
Job posting may be removed earlier if the position is filled or if a sufficient number of applications are received.

This is a Hybrid position requiring approximately 2 days per week on-site at Cisco offices in either San Francisco, San Jose, or New York City.

The Splunk Agent Resilience team is defining the future of AI resilience. Together our team provides scalable, cost-effective evaluation and guardrails that ensure AI agents behave as intended, improving reliability and reducing risks. This unified approach empowers our customers to expertly deploy and manage AI-powered applications with enhanced observability and control.

As a Staff Site Reliability Engineer (SRE), you will provide technical leadership for the reliability, scalability, and operational architecture of Splunk Agent Observability's platform. You will define the long-term reliability strategy, lead major infrastructure initiatives, and drive engineering excellence across deployment automation, production operations, and platform resiliency. In addition to owning complex production systems, you will influence engineering direction across teams and establish best practices for operating large-scale cloud and on-prem deployments.

Your Impact
  • Define and drive the technical roadmap for platform reliability, scalability, and operational excellence.
  • Lead the architecture and evolution of deployment platforms supporting cloud and air-gapped customer environments.
  • Establish reliability engineering standards, including service level objectives (SLOs), operational readiness, capacity planning, and resiliency reviews.
  • Lead major reliability and scalability initiatives across Kubernetes, deployment infrastructure, databases, and networking.
  • Drive automation to eliminate operational toil and improve engineering productivity.
  • Design and build internal platforms, frameworks, and tooling that enable reliable operations at scale.
  • Lead complex production incident response and drive systemic improvements through root cause analysis and long-term remediation.
  • Partner with engineering leadership to influence platform architecture, deployment strategy, and production readiness.
  • Mentor engineers and raise the engineering bar through technical leadership, design reviews, and operational guidelines.
  • Collaborate with customers and internal teams to build secure, scalable, and highly reliable deployment architectures for both cloud and on-prem environments.


Minimum Qualifications
  • 8+ years' experience with a Bachelor's degree or 6+ yrs with Masters or 3+ years with a PhD, or equivalent related experience; Experience to include at least 6 years in Site Reliability Engineering, Platform / Cloud / Infrastructure Engineering, or related fields.
  • 5+ years' operating large-scale Kubernetes platforms in production.
  • Experience designing highly available, scalable, and resilient distributed systems.
  • Experience with AWS, GCP, or other public cloud platforms.
  • Strong experience designing CI/CD platforms and deployment automation at scale.


Preferred Qualifications
  • Expertise in observability, monitoring, alerting, capacity planning, and performance engineering.
  • Strong programming skills in Python and/or Go
  • Deep experience with Infrastructure as Code (Terraform or similar)
  • Experience with MLOps preferred
  • Strong understanding of networking, distributed systems, storage, databases, and cloud architecture
  • Experience operating both SaaS and enterprise/on-prem deployments
  • Demonstrated technical leadership across multiple engineering teams
  • Experience leading incident management, postmortems, and long-term reliability initiatives

About Cisco

Cisco Careers

Join the vibrant team at Cisco, a global leader in networking and cybersecurity solutions, where innovation and leadership thrive. Cisco offers a plethora of job opportunities that cater to a range of skills and experiences, making it an ideal place for both seasoned professionals and those seeking an internship to jumpstart their career. Work You’ll Do At Cisco, you’ll be part of a culture that values diversity, leadership, and professional growth. Engage in work that matters with a team that combines technology, creativity, and the power of human connection to redefine networking. Cisco’s commitment to innovation isn’t just about technology, but also about transforming the way we work and collaborate. Cisco’s employment philosophy supports career advancement and nurtures a leadership pipeline that is equipped with diversity training and opportunities for growth. Whether you’re applying your skills to drive our latest innovations or using our vast networking capabilities to solve complex problems, at Cisco, every role is impactful. Join Our Dynamic Team Explore job opportunities in areas ranging from engineering to marketing, sales to cybersecurity. Cisco is hiring individuals who are passionate, curious, and ready to drive change. Positions at Cisco offer competitive benefits, a supportive culture, and the chance to work with cutting-edge technology. Internship Programs Kickstart your career with a Cisco internship. Gain invaluable industry experience, enhance your resume, and build professional networks that last a lifetime. Our internships provide hands-on experience and the chance to work on projects that matter. Leadership and Development Cisco is committed to fostering leadership skills and providing employees with the training needed to succeed. Our leadership programs help you develop new skills, manage teams effectively, and lead with confidence. Cisco’s commitment to professional development ensures that your career path is as dynamic as our technologies. Benefits and Culture Cisco understands the importance of a balanced life. Our benefits package is designed to ensure that our team members are healthy, happy, and secure. At Cisco, you’ll find a supportive culture that encourages open communication, teamwork, and mutual respect. Stay Connected Join Cisco’s Talent Network Stay informed about new positions that match your skills and interests. At Cisco, we value the curiosity and unique perspectives of our team members. Subscribe to receive personalized job alerts and insider tips directly from our hiring managers. Explore Cisco Jobs Ready to advance your career at Cisco? Search open positions, prepare your resume, and get ready for an interview that could lead to your next big opportunity. At Cisco, we’re not just filling positions—we’re investing in leaders. Keep Up to Date Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here. READ CAREERS BLOG Job Alert Emails Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await you at Cisco.
Learn more about Cisco
Size
79,500 employees
Market Cap
$194.5 billion
Industry
Net Income
$10.1 billion
Founded
2014
5 Year Trend
+1.4%
Revenue
$48 billion
NASDAQ

Similar Jobs

More Jobs at Cisco

More Information Technology Jobs

Find similar Staff Site Reliability Engineer (SRE) (Hybrid) jobs: