To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.
Job Category
Enterprise Technology & Infrastructure
Job Details
As a Site Reliability Operations Engineer you'll be part of our internal DET Site Reliability Operations team supporting our employees globally. This role combines incident command, reliability engineering, and hands-on technical support. You'll help keep critical systems running while working with teams across different time zones.
What You'll Actually Be Doing...- Respond to and manage major incidents affecting internal business operations. Serve as Incident Commander to coordinate technical teams, establish impact, and drive rapid service restoration.
- Monitor and troubleshoot enterprise systems including infrastructure, applications, and network components. Use your technical skills to diagnose complex problems across multiple platforms and vendors before they impact users.
- Work with teams globally to improve incident response by creating and improving runbooks, developing SOPs, and driving automation.
- Coordinate emergency changes and infrastructure updates to resolve incidents. Work with cross-functional teams to maintain business continuity during critical situations.
- Analyze incident data and KPI metrics to identify trends. Develop actionable recommendations to reduce impact duration and improve performance, then present findings to stakeholders.
- Lead problem management activities, investigating recurring incidents, documenting root cause analyses, and tracking known errors.
- Participate in on-call rotation as part of regional coverage. Handle escalations during your shift and serve as Duty Manager for high severity incidents when needed.
- Track on-call burden and surface toil reduction opportunities with measurable impact.
You're Our Person If...- 5-8 years in IT operations, incident management, or site reliability work. Experience in a 24x7 high availability environment with enterprise systems preferred.
- Demonstrated ability to manage high severity incidents under pressure. Establish impact, evaluate solutions with subject matter experts, and make decisions that balance technical and business needs.
- Strong verbal and written communication skills to explain complex technical issues to both technical and executive audiences. Create clear incident updates and status reports.
- Demonstrated technical troubleshooting ability across Windows and Linux servers, networking, cloud platforms, and virtualization technologies. Diagnose problems quickly using logs, monitoring tools, and common diagnostic approaches.
- Experience with cloud platforms (e.g. AWS) and monitoring of IT infrastructure. You should know core cloud concepts and be comfortable with monitoring tools.
- Understanding of ITIL framework, particularly incident, problem, and change management processes.
- A related technical degree required.
Even Better If...- Salesforce platform experience and certifications
- Industry certifications like ITIL, AWS, CCNA, MCSA, or RHCE
- Scripting ability in Python, Bash, PowerShell, or similar languages to help automate, reduce manual work, and improve efficiency.
- Experience with monitoring and visualization tools like Splunk, Grafana, or Tableau. Ability to analyze data and identify trends for improving reliability.
- Background with automation tools like Puppet or Chef
Unleash Your Potential
When you join Salesforce, you'll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best, and our AI agents accelerate your impact so you can do your best. Together, we'll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future - but to redefine what's possible - for yourself, for AI, and the world.
Accommodations
If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form.
Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates' resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.
At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions.The typical base salary range for this position is $94,000 - $142,300 annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.