Cox Communications

Sr Software Engineer - Reliability Engineering

Cox Communications$121K — $203K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in software, platform, or infrastructure engineering.
  • Proficient in coding languages such as Python, Go, or Java.
  • Hands-on experience with AWS services including EC2 and Lambda.
  • Skilled in Terraform or similar infrastructure-as-code tools.
  • Familiarity with Docker and Kubernetes for container orchestration.
  • Windows and Linux troubleshooting capabilities underlined by comprehensive system debugging skills.
  • Bachelor's degree or an equivalent combination of education and experience.

Responsibilities

  • Design and implement resilience features like failover and disaster recovery.
  • Write and manage infrastructure-as-code in Terraform across multiple AWS accounts.
  • Proactively maintain system health and minimize downtime.
  • Build observability for applications including logs and metrics.
  • Improve incident response through enhanced monitoring frameworks.
  • Drive AWS cost optimization through architectural improvements.
  • Develop production systems and mentor junior engineers in best practices.

Benefits

  • Unlimited paid vacation policy based on duty and company needs.
  • Seven paid holidays per year.
  • Up to 160 hours of paid wellness leave for personal or family wellness.
  • Additional leave options for bereavement, voting, jury duty, volunteering, military service, and parental reasons.
Full Job Description
Job Family Group

Engineering / Product Development

Job Profile

Sr Software Engineer

Management Level

Individual Contributor

Flexible Work Option

Hybrid - Ability to work remotely part of the week

Travel %

Yes, 5% of the time

Work Shift

Day

Compensation
Compensation includes a base salary in the range of $121,800.00 - $203,000.00. The base salary may vary within the anticipated base pay range based on factors such as the ultimate location of the position and the selected candidate's knowledge, skills, and abilities. Position may be eligible for additional compensation that may include an incentive program.

Job Description

We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems-treating resilience and debuggability as first-class concerns. You'll own projects end-to-end: from architecture 12 code 12 deployment 12 production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily.

What You'll Do:

SRE Best Practices & System Health Management

  • Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
  • Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
  • Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.
  • Evolve team's SRE standards and practices.


Application Monitoring & Observability

  • Build observability into systems: logging, metrics, distributed tracing, alert design.
  • Improve monitoring frameworks; enable faster incident detection and resolution.
  • Design dashboards and alerts that help teams understand system behavior.
  • Partner with application teams on service instrumentation.


AWS Cost Optimization

  • Drive significant reductions in cloud spend through architectural improvements and resource utilization.
  • Review infrastructure for efficiency; identify and eliminate waste.
  • Balance cost, performance, and reliability in design decisions.


Software Development & Architecture

  • Build production systems, APIs, internal tools, and automation with clean, well-tested code.
  • Design for maintainability, operational simplicity, and reliability.
  • Participate in code review and technical design discussions.
  • Mentor junior engineers on code quality and architectural thinking.


Operations & Incident Response

  • Participate in on-call rotations; debug and resolve production incidents.
  • Conduct postmortem analysis; drive systemic improvements.
  • Develop operational procedures and runbooks.


Qualifications:
  • 5+ years software engineering, platform engineering, or infrastructure engineering experience.
  • Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
  • AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
  • Terraform or equivalent infrastructure-as-code experience.
  • Docker and container orchestration (Kubernetes or similar).
  • Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
  • System design thinking: can architect scalable systems and reason about trade-offs.
  • Bachelor's degree in a related discipline and 4 years' experience in a related field. The right candidate could also have a different combination, such as a master's degree and 2 years' experience; a Ph.D. and up to 1 year of experience; or 16 years' experience in a related field.


Highly Valued:

  • Experience with observability tools (New Relic, Splunk, Prometheus).
  • Incident response experience; familiar with postmortem practices.
  • Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).
  • Cost optimization mindset; has identified and eliminated cloud waste.
  • Windows and Linux system troubleshooting and performance analysis.
  • Experience with CI/CD pipelines and deployment automation.


Drug Testing

To be employed in this role, you'll need to clear a pre-employment drug test. Cox Automotive does not currently administer a pre-employment drug test for marijuana for this position. However, we are a drug-free workplace, so the possession, use or being under the influence of drugs illegal under federal or state law during work hours, on company property and/or in company vehicles is prohibited.

Benefits

The Company offers eligible employees the flexibility to take as much vacation with pay as they deem consistent with their duties, the company's needs, and its obligations; seven paid holidays throughout the calendar year; and up to 160 hours of paid wellness annually for their own wellness or that of family members. Employees are also eligible for additional paid time off in the form of bereavement leave, time off to vote, jury duty leave, volunteer time off, military leave, and parental leave.

About Cox Communications

Cox Communications is a telecommunications company that provides cable television, internet, and phone services to residential and business customers. The company operates in 18 states and has over 6 million customers. Cox Communications is a subsidiary of Cox Enterprises, a privately held company that also owns newspapers, television stations, and radio stations. The company was founded in 1962 and is headquartered in Atlanta, GA.
Learn more about Cox Communications
Size
20,000 employees
Industry
Net Income
$2 billion
5 Year Trend
+10%
Revenue
$12 billion
NASDAQ

Similar Jobs

More Jobs at Cox Communications

More Information Technology Jobs

Find similar Sr Software Engineer - Reliability Engineering jobs: