Cox Communications

Sr Software Engineer - Reliability Engineering

Cox Communications$121K — $203K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in software engineering or related fields.
  • Strong coding skills in Python, Go, Java, or similar languages.
  • Hands-on experience with AWS services (EC2, RDS, S3, etc.).
  • Proficient in Terraform or equivalent infrastructure-as-code technologies.
  • Experience with Docker and container orchestration (Kubernetes).
  • Debugging skills on Linux and Windows; adept at troubleshooting complex systems.
  • Bachelor's degree in a related discipline or equivalent experience.

Responsibilities

  • Design and implement resilience initiatives like redundancy and disaster recovery.
  • Write infrastructure-as-code using Terraform; manage extensive AWS accounts.
  • Maintain application performance and user experience proactively.
  • Build observability into systems including logging and metrics.
  • Drive cloud cost reduction through architectural improvements.
  • Develop production systems and internal automation tools with quality code.
  • Participate in on-call support and conduct postmortem analyses.

Benefits

  • Unlimited vacation based on individual discretion and company needs.
  • Seven paid holidays each year.
  • Up to 160 hours of paid wellness leave annually.
  • Additional paid time off for bereavement, voting, and jury duty.
  • Flexible parental and military leave options.
Full Job Description
Sr Software Engineer - Reliability Engineering Cox Automotive - USA Engineering / Product Development Sr Software Engineer Individual Contributor Hybrid - Ability to work remotely part of the week Yes, 5% of the time Day Compensation includes a base salary in the range of $121,800.00 - $203,000.00. The base salary may vary within the anticipated base pay range based on factors such as the ultimate location of the position and the selected candidate’s knowledge, skills, and abilities. Position may be eligible for additional compensation that may include an incentive program. We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems—treating resilience and debuggability as first-class concerns. You'll own projects end-to-end: from architecture  code  deployment  production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily. What You'll Do: SRE Best Practices & System Health Management  Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.  Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.  Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.  Evolve team's SRE standards and practices. Application Monitoring & Observability  Build observability into systems: logging, metrics, distributed tracing, alert design.  Improve monitoring frameworks; enable faster incident detection and resolution.  Design dashboards and alerts that help teams understand system behavior.  Partner with application teams on service instrumentation. AWS Cost Optimization  Drive significant reductions in cloud spend through architectural improvements and resource utilization.  Review infrastructure for efficiency; identify and eliminate waste.  Balance cost, performance, and reliability in design decisions. Software Development & Architecture  Build production systems, APIs, internal tools, and automation with clean, well-tested code.  Design for maintainability, operational simplicity, and reliability.  Participate in code review and technical design discussions.  Mentor junior engineers on code quality and architectural thinking. Operations & Incident Response  Participate in on-call rotations; debug and resolve production incidents.  Conduct postmortem analysis; drive systemic improvements.  Develop operational procedures and runbooks. Qualifications:  5+ years software engineering, platform engineering, or infrastructure engineering experience.  Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.  AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.  Terraform or equivalent infrastructure-as-code experience.  Docker and container orchestration (Kubernetes or similar).  Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.  System design thinking: can architect scalable systems and reason about trade-offs.  Availability for rotational on-call duties outside of standard business hours may be required.  Bachelor’s degree in a related discipline and 4 years’ experience in a related field. The rightcandidate could also have a different combination, such as a master’s degree and 2 yearsexperience; a Ph.D. and up to 1 year of experience; or 16 years’ experience in a related field. Highly Valued:  Experience with observability tools (New Relic, Splunk, Prometheus).  Incident response experience; familiar with postmortem practices.  Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).  Cost optimization mindset; has identified and eliminated cloud waste.  Windows and Linux system troubleshooting and performance analysis.  Experience with CI/CD pipelines and deployment automation. Drug Testing To be employed in this role, you’ll need to clear a pre-employment drug test. Cox Automotive does not currently administer a pre-employment drug test for marijuana for this position. However, we are a drug-free workplace, so the possession, use or being under the influence of drugs illegal under federal or state law during work hours, on company property and/or in company vehicles is prohibited. Benefits The Company offers eligible employees the flexibility to take as much vacation with pay as they deem consistent with their duties, the company’s needs, and its obligations; seven paid holidays throughout the calendar year; and up to 160 hours of paid wellness annually for their own wellness or that of family members. Employees are also eligible for additional paid time off in the form of bereavement leave, time off to vote, jury duty leave, volunteer time off, military leave, and parental leave.

About Cox Communications

Cox Communications is a telecommunications company that provides cable television, internet, and phone services to residential and business customers. The company operates in 18 states and has over 6 million customers. Cox Communications is a subsidiary of Cox Enterprises, a privately held company that also owns newspapers, television stations, and radio stations. The company was founded in 1962 and is headquartered in Atlanta, GA.
Learn more about Cox Communications
Size
20,000 employees
Industry
Net Income
$2 billion
5 Year Trend
+10%
Revenue
$12 billion
NASDAQ

Similar Jobs

More Jobs at Cox Communications

More Information Technology Jobs

Find similar Sr Software Engineer - Reliability Engineering jobs: