Job DescriptionWe are looking for a talented
Sr. Site Reliability Engineer (SRE) to join our Observability team. In this role, you will work alongside other engineers to help build, maintain, and improve reliable, secure, and scalable observability solutions for our global infrastructure.
As an Sr. SRE, you will help bridge the gap between development and operations, ensuring our systems and services are robust and well-monitored. You will assist in automating workflows, deploying and maintaining monitoring tools, and supporting the smooth operation of our technology platforms.
You will participate in troubleshooting issues, responding to incidents, and implementing fixes under the guidance of senior team members. You'll also contribute to building and enhancing automation, documentation, and reliability practices
Responsibilities- Design, deploy, and enhance monitoring and logging instrumentation to ensure comprehensive observability.
- Develop and maintain automation to streamline operational workflows and improve system integration.
- Operate, support, and optimize observability platforms and tools, including Splunk, ClickHouse, Grafana, Prometheus, M3DB, OpenTelemetry, Fluent Bit, ElasticSearch, OpenSearch, and CloudWatch.
- Collaborate with development teams to architect monitoring solutions and assist in building development, staging, and production environments.
- Safeguard systems by proactively applying security hotfixes and OS patches to mitigate cybersecurity threats.
- Champion DevOps best practices across all platforms and teams.
- Design, implement, and maintain CI CD pipelines to automate build, test, and deployment processes.
- Manage and optimize cloud infrastructure (AWS, GCP) to ensure availability, scalability, performance, and security.
- Leverage infrastructure-as-code tools such as Terraform, Ansible, or CloudFormation to automate infrastructure provisioning and management.
- Monitor system health, diagnose issues, and deliver solutions to enhance reliability and efficiency.
- Implement and manage containerization and orchestration platforms, such as Docker and Kubernetes, for streamlined application deployment.
- Conduct root cause analysis for production incidents and implement preventative solutions.
- Maintain clear and up-to-date documentation for infrastructure, procedures, and processes.
- Provide technical support and troubleshooting for infrastructure and deployment-related issues.
- Continuously develop automation to boost operational efficiency and improve system integration.
This is a hybrid position. Expectation of days in office will be confirmed by your Hiring Manager.
Visa requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager.
QualificationsBasic Qualifications
• Bachelor's degree with 3+ years of relevant professional experience, OR Advanced degree (e.g., Master's, MBA, JD, MD) with at least 2 years of relevant experience, OR 6+ years of relevant industry experience
Preferred Qualifications
• Hands-on experience performing upgrades and maintenance for observability tools such as Splunk Enterprise and Universal Forwarder, Fluent Bit, Prometheus container images, Grafana, ElasticSearch, OpenSearch, Kibana, and ClickHouse.
• Experience configuring and deploying exporters (e.g., Node Exporter, Cert Exporter) for metrics collection.
• Proficient with containerization and orchestration platforms, including Docker and Kubernetes.
• Strong background working with observability solutions such as Splunk, ClickHouse, Grafana, Prometheus, M3DB, OpenTelemetry, Fluent Bit, ElasticSearch, OpenSearch, and CloudWatch.
• Experience managing CI CD pipelines using tools like GitHub and Ansible.
• Familiarity with Infrastructure as Code tools (e.g., Terraform) and configuration management approaches such as GitOps.
• Scripting abilities in Python and or Shell.
• Working knowledge of query languages such as PromQL, MS SQL, or Splunk SPL is a plus.
• Solid experience in Linux environments and Unix shell scripting.
• GCP or AWS cloud certifications are a plus.
• Strong analytical skills with the ability to assess problems and solutions at the appropriate level of detail to engage with various stakeholders.
• Excellent communication and leadership abilities.
Justification
Visa's Observability ecosystem includes over 2,000 platform nodes, utilizing approximately 15 different tools for logging, monitoring, and tracing, alongside 80,000 client agents. The system handles daily log ingestion exceeding 100TB and oversees hundreds of critical applications, supporting vital alerts, dashboards, and reports. To maintain this high level of performance and reliability, we need a Sr. Site Reliability Engineer (SRE) with comprehensive knowledge and practical experience. This position requires an I5-level engineer who can operate independently with minimal supervision.
About Visa's PRE Observability Team
Visa's Product Reliability Engineering (PRE) Observability team partners with Product Development as well as Operations & Infrastructure teams to build and manage innovative, reliable, scalable, secure, and cost-effective observability platform solutions. We are looking for talented Senior Site Reliability Engineers to join our driven team, with a focus on maximizing system availability, performance, security, and reliability. This dynamic role requires technical leadership, strong problem-solving skills, and expertise in coding, testing, and debugging.Information for US Applicants :For roles located in the US, the estimated salary range for this position is $110,700.00 to $ 171,800.00 USD per year, which may include potential sales incentive payments (if applicable). Salary may vary depending on job-related factors which may include knowledge, skills, experience, and location. In addition, this position may be eligible for bonus and equity.Visa has a comprehensive benefits package for which this position may be eligible that includes Medical, Dental, Vision, 401(k), FSA/HSA, Life Insurance, Paid Time Off, and Wellness Program.
Work HoursVaries upon the needs of the department.
Travel RequirementsThis position requires travel 5-10% of the time.
Mental/Physical RequirementsThis position will be performed in an office setting. The position will require the incumbent to sit and stand at a desk, communicate in person and by telephone, frequently operate standard office equipment, such as telephones and computers.