JobOverview:
The Senior Engineer, Observability and Platform Stability is responsible for ensuring the reliability, availability, and performance of enterprise observability platforms and supporting applications. This role drives operational excellence through proactive monitoring, incident response, platform maintenance, automation, and continuous improvement initiatives. The position partners closely with Site Reliability Engineering (SRE), Platform Engineering, DevOps, Operations, and Incident Management teams to enhance platform stability, mature CI/CD practices, and support strategic technology initiatives. The ideal candidatebrings strong cloud technology experience and a proven ability to support production environments while delivering actionable insights through observability capabilities.
Responsibilities:
DeliverySupport
Partnerwith Product Owners and engineering teams to provide operational and release support for technology initiatives.
Ensureplatform changes meet operational readiness requirements, including rollback procedures, runbook documentation, integration standards, and support handoffs.
Maintainproduction stability throughout platform upgrades, enhancements, and enterprise initiatives.
Supportplatform ownership transitions and operational readiness activities across global delivery teams.
ProductionSupport & Incident Response
Troubleshootapplication and platform issues to restore services and minimize business impact.
Serveas an escalation point for complex production incidents and operational challenges.
Participatein incident triage, root cause analysis, corrective action planning, and resolution activities.
Collaboratewith engineering teams to implement long-term solutions that reduce recurring incidents.
Supportmission-critical environments through timely incident response and service restoration.
PlatformStability & Proactive Operations
Conducthealth checks, configuration reviews, and performance assessments to identify operational risks.
Supportreliability initiatives focused on improving recovery times, reducing incident recurrence, and minimizing change-related defects.
Validatevendor releases, hotfixes, and configuration changes prior to production deployment.
Partnerwith observability and analytics teams to enhance monitoring, alerting, and issue detection capabilities.
ReleaseExecution & CI/CD Maturation
Executeplatform changes through established SDLC, change management, and release management processes.
Collaboratewith Platform Engineering, DevOps, Quality Engineering, and Scrum teams to improve release and deployment practices.
Supportthe adoption of source control, environment separation, release automation, and CI/CD capabilities.
Ensuresolutions are testable, deployable, and operationally supported before and after production implementation.
Documentation& Operational Excellence
Maintainrunbooks, support documentation, configuration records, incident playbooks, and release procedures.
Documentincident findings, lessons learned, and process improvement opportunities.
Contributeto the development of standardized, repeatable, and scalable operational practices.
Whatare we looking for?
Weseek professionals who pursue greatness, act with integrity, are driven to help our clients succeed, win together, and create and share joy. The ideal candidate brings strong technical expertise in observability platforms, site reliability engineering, production support, release management, and operational excellence while demonstrating a commitment to platform reliability, cross-functional collaboration, and continuous improvement.
Requirements:
Bachelor’sdegree in Computer Science, Information Technology, Engineering, or a related field.
6+years of experience in SRE, platform engineering, production support, enterprise application operations, or related technology environments.
5+years of experience supporting enterprise platforms, including AWS, Dynatrace, ELK, ServiceNow, and SolarWinds.
Experiencetroubleshooting complex production incidents within enterprise-scale technology environments.
Experienceexecuting technology changes through formal change management and release management processes.
Preferences:
Experiencewithin financial services or another regulated industry.
Experienceimplementing or supporting CI/CD pipelines, release automation, and deployment processes across development, testing, and production environments.
Experiencewith observability platforms, monitoring tools, performance dashboards, or application monitoring solutions.
Experiencecollaborating with offshore, nearshore, or global delivery teams.
Pay Range:
$101,558.00 - $169,229.00
Actual base salary varies based on factors, including but not limited to, relevant skill, prior experience, education, base salary of internal peers, demonstrated performance, and geographic location. Additionally, LPL Total Rewards package is highly competitive, designed to support your success at work, at home, and at play – such as 401K matching, health benefits, employee stock options, paid time off, volunteer time off, and more. Your recruiter will be happy to discuss all that LPL has to offer!