Job Summary
We are seeking a Secure Enterprise Browser DevOps Engineer with production support experience to ensure the reliability, availability, and performance of a business-critical secure enterprise browser platform. The role will focus on software-driven operations, production reliability, access and identity controls, browser and extension lifecycle management, automation, observability, and incident response. The selected candidate will share production ownership and on-call responsibilities with SRE and incident response teams while helping eliminate operational toil and strengthen the platform against large-scale incidents.
Key Responsibilities
• Engineer, test, and deploy access and identity controls, including group policies and SAML-based roles.
• Manage browser version and extension lifecycles through safe change processes, staged rollouts, validation, and rapid rollback.
• Build automation for deployment, configuration, monitoring, health checks, and other operational activities.
• Define service-level objectives and implement observability, performance analytics, and capacity monitoring to maintain platform health at scale.
• Own production reliability and participate in on-call responsibilities for business-critical systems.
• Lead incident response during outages and mass-impact events.
• Perform advanced troubleshooting and tier-4 technical support for complex production issues.
• Conduct root-cause analysis and lead blameless post-incident reviews.
• Translate incident findings into permanent fixes and reliability improvements.
• Develop and maintain infrastructure-as-code and configuration management solutions.
• Automate repetitive operational processes to reduce manual effort and improve platform reliability.
• Implement monitoring and alerting to proactively identify system health, performance, and availability issues.
• Maintain secure access controls and enterprise browser configurations in accordance with organizational requirements.
• Support browser platform scalability, performance, and operational readiness as the environment grows.
Required Qualifications
• Experience operating business-critical production systems with ownership of reliability, availability, and on-call responsibilities.
• Strong software-driven operations experience using Linux administration and coding with Python and Bash.
• Experience with infrastructure-as-code and configuration management.
• Experience implementing endpoint or enterprise browser controls, including group policies, SAML, SSO, and extension governance.
• Experience building and maintaining CI/CD pipelines.
• Experience implementing observability, monitoring, and automated health checks to reduce operational toil.
• Experience leading incident response and root-cause analysis for production environments.
• Ability to translate incident findings into lasting reliability and operational improvements.
• Strong troubleshooting and problem-solving skills in complex production environments.
Preferred Qualifications
• Experience with Site Reliability Engineering practices, including SLOs, SLIs, error budgets, and blameless postmortems.
• Experience with enterprise browser, VDI, or thin-client platforms.
• Experience implementing progressive delivery patterns such as canary deployments and feature-flagged rollouts.
• Experience implementing automated rollback capabilities.
• Experience delivering solutions in regulated environments with audit logging and operational controls.
• Experience using SQL for telemetry analytics and operational dashboards.
• Experience designing reliability, availability, and performance improvements for large-scale enterprise platforms.
• Strong understanding of production operations, automation, monitoring, and incident management.