JOB DESCRIPTION
DESCRIPTION:
Duties: Define and enforce measurable reliability targets for critical application environments, ensuring operational metrics are consistently tracked and achieved. Architect, implement, and refine advanced monitoring and alerting solutions to proactively surface application health and performance issues across complex technology stacks. Collaborate with cross-functional teams to design and support robust, highly available application platforms capable of meeting stringent business and technical demands. Drive adoption of reliability and resilience best practices providing technical leadership to elevate operational standards across support and engineering groups. Develop, automate, and maintain failover and recovery workflows for application services ensuring uninterrupted operations across diverse infrastructure and cloud regions. Create, update, and operationalize incident response documentation and automated remediation mechanisms including rollback and circuit breaker strategies for application failures. Lead critical incident management for application platforms ensuring rapid restoration, comprehensive root cause analysis, and implementation of long-term improvements. Optimize resource allocation and infrastructure costs for large-scale application support, balancing efficiency with reliability, and performance requirements. Facilitate seamless integration and deployment of application changes bridging development and support to ensure smooth operational transitions. Oversee continuous validation processes including pre- and post-deployment monitoring to detect and remediate application drift and performance regressions. Maintain day-to-day operational stability and high availability for application systems, leveraging deep technical expertise in support and troubleshooting. Monitor production environments using advanced diagnostic and observability tools rapidly identifying and resolving anomalies. Escalate and communicate complex technical issues, delivering actionable insights and solutions to both technical and business stakeholders.
QUALIFICATIONS:
Minimum education and experience required: Bachelor's degree in Information Systems Engineering, Computer Engineering, or related field of study plus 5 years of experience in the job offered or as Site Reliability Engineer, Data Engineer, Data Analyst, MSSQL Server Developer, Support Engineer, Software Developer, or related occupation.
Skills Required: This position requires five (5) years of experience with the following: Utilizing Sev1/Sev2 on-call including triaging, mitigating, coordinating restoration, RCA, and blameless postmortems; MTTR recurrence reduction; Observability including metrics, logs, traces; low-noise alerts, fast detection; maintaining runbooks and escalations; Automation including Python and Bash; utilizing CI/CD with blue and green or canary, feature flags, pre-deploy validation, and reliable rollback; using IaC including Terraform for public cloud; using Modules, remote state, drift detection, LUT/policy-as-code, and targeted applies; using Kubernetes for deployments, autoscaling, health probes, progressive rollouts and rollbacks, quotas, and cluster troubleshooting; utilizing AWS production operations including VPC networking, IAM policy design, encryption/ KMS, load balancing, multi-AZ resilience, backup and restore, and regional failover; Database reliability including PostgreSQL, MySQL, Oracle backups/PITR, replication and failover, online schema changes, SQL and index tuning under load; Linux and networking including processing memory/IO diagnostics, kernel and sysctl tuning, TLS, TCP/IP, DNS, HTTP, load balancing, and service discovery; Security including least privilege, secrets management, patch, vulnerability management, immutable audit logging and framework alignment. This position requires three (3) years of experience with the following: Utilizing SLI/SLO and error budgets sustaining 69.9% SLOs; Gate high-risk changes; Performance and capacity including load testing (JMeter, BlazeMeter), trace hot paths, tail-latency/throughput tuning, and peak demand modeling; Resilience and DR including timeouts, retries, backpressure, circuit breaking, graceful degradation; validated RTO/RPO; using Linux and networking to process memory/IO diagnostics, kernel and sysctl tuning, TLS, TCP/IP, DNS, HTTP, load balancing, and service discovery; Releasing change management including versioned applications, DB migrations, automated quality gates Maxwell, controlled rollouts, and rapid clean NB rollback; Reliability outcomes including sustaining lower bast MT pipeline failure TR, fewer false alarms, higher S, logical improvements, safer deployments, hardened failure modes, and predictable peak scaling.
Job Location: 8181 Communications Pkwy, Plano, TX 75024.
Full-Time.