Geico

Senior Engineer

Geico • $100K — $215K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Hands-on proficiency in Go, Java, Python, or C# for building applications on Kubernetes and serverless technologies.
  • Experience with SQL and NoSQL for incident data management.
  • Familiarity with data pipelines and analytics for operational metrics using tools like Spark and Grafana.
  • Knowledge of observability platforms such as Grafana and Datadog for incident monitoring.
  • Experience with incident management tools like PagerDuty and automation processes.

Responsibilities

  • Design and develop automation and self-service tools for incident management.
  • Build shared services and APIs to standardize incident response.
  • Apply engineering standards for deployment, testing, and support.
  • Implement CI/CD and operational controls for reliable delivery.
  • Act as a technical leader during high-severity incidents.

Benefits

  • Collaborative work culture emphasizing continuous improvement and accountability.
  • Exposure to cutting-edge technologies and cloud services.
  • Opportunity to influence critical incident management processes.
  • Participation in shaping engineering excellence and operational practices.
Full Job Description
Position Summary

Technology Operations Center is at the core of GEICO's application and platform resiliency. It assists GEICO's engineering teams with maintaining high availability of our customer and internally facing services while driving down time to detect and recover from incidents. It is governing our incident management processes and builds platforms that allow GEICO to manage and recover from incidents.

GEICO is seeking an experienced SRE Software Engineer with a passion for building, operating and troubleshooting high-performance, low-maintenance, zero-downtime complex distributed platforms and applications. You will help drive our transformation to a tech organization with engineering excellence and site reliability as its mission, by implementing and operating our incident management processes and in-house technology platforms that automate them.

This role focuses on improving Incident Management tooling and process across GEICO. It is a hands-on engineering role focused on better incident management, faster time to detect, troubleshoot and recover from incidents and fewer repeat incidents, by designing, developing and operating the tools and processes that help GEICO engineering teams manage their on-call, runbooks, troubleshooting and BCDR.

Success in this role requires strong technical depth and effective process execution. The right candidate can go deep on incident analysis, system behavior, and design, while also helping teams improve COEs, learn from incidents, and turn those lessons into engineering improvements.

Why This Role Is Different
  • This role blends deep technical understanding, hands-on execution and the ability to build software with real-time incident leadership, platform and process improvements. You will help drive our incident management and response processes and the platforms that automate them.
  • You are helping shape how GEICO service engineering teams detect, manage, resolve and learn from incidents.
  • Your work directly impacts the availability of GEICO's critical applications and platforms, experiences and satisfaction of millions of customers and tens of thousands of associates.

Position Responsibilities

As a Senior Engineer in the TOC, you will:

Build and Operate Enterprise-Critical Platforms
  • Design, develop and operate automation, self-service tools, dashboards, and data pipelines that automate and scale our incident management, on-call, paging and troubleshooting processes.
  • Build shared services, APIs, data contracts, automation, and integrations that standardize incident response and reduce operational risk.
  • Apply engineering standards across design, implementation, deployment, testing, observability, security, operational support, and production readiness.
  • Implement safe deployment, CI/CD, infrastructure as code, automated testing, rollback patterns, and operational controls that support frequent and reliable delivery.
  • Evaluate and implement modern technologies and tools that improve platform capability, compliance, visibility, reliability, and engineering effectiveness.

Serve as a Technical Leader During Incidents
  • Act as a technical leader during high-severity incidents, bringing sound technical judgment, system-level problem solving, and calm execution under pressure.
  • Contribute to troubleshooting strategy, cross-team coordination, impact analysis, and risk-based decision making to restore service safely and efficiently.
  • Contribute substantially to post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements.
  • Develop and maintain operational runbooks, readiness criteria, triage models, and resilience practices across assigned integration points.

Contribute to Technical Direction and Engineering Culture
  • Contribute to design and architecture reviews spanning teams, services, dependencies, and operational domains.
  • Partner with SRE, platform, product, infrastructure, security, and business stakeholders to align operational tooling with practical engineering needs and enterprise reliability goals.
  • Translate technical concepts, risks, and tradeoffs clearly for technical leaders and non-technical stakeholders.
  • Mentor engineers through technical leadership, example, code and design reviews, documentation, and operational coaching.
  • Reinforce a culture of ownership, accountability, continuous improvement, psychological safety, learning, and operational excellence.

We have adopted a "You Build It, You Run It" strategy.
  • All our senior technologists take an active role in leading and managing high-severity incidents requiring strong technical judgment, clear communication, and calm execution under pressure.
  • All our engineers have on-call responsibilities as part of a 24x7 rotation supporting incident response and production support for mission-critical platforms and processes they build and operate.

Qualifications
  • Hands-on proficiency in multiple languages, including Go, Java, Python, or C#, for building production-grade full stack applications on Kubernetes and serverless technologies such as KNative in Azure and AWS.
  • Experience with SQL and NoSQL technologies and cloud-native services for storing and analyzing incident data.
  • Experience with building and using data pipelines, analytics, and dashboards for operational metrics, trends, and KPIs using technologies such as Spark, Trino, Grafana, Superset, or Power BI.
  • Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, or Azure Monitor.
  • Experience with incident management platforms such as PagerDuty.
  • Proficiency with AI-assisted development processes and tools such as Claude Code, Cursor, and GitHub Copilot.
  • Experience improving incident, post-incident review, or reliability processes through automation, data, and cross-team collaboration.
  • Strong incident forensics and root cause analysis skills, with the ability to improve COE quality through clear action items and follow-through.
  • Strong understanding of observability, reliability engineering, incident management, and post-incident improvement practices.
  • Experience supporting incident response and high-severity production incidents in complex environments.
  • Strong software engineering fundamentals and system design skills, with experience building reliable production systems.
  • Ability to contribute to technical design and architecture decisions in complex distributed systems.
  • Strong communication skills and the ability to support engineering teams and present findings clearly to leadership.

Experience
  • 4+ years of professional software engineering experience, preferably in platform engineering, reliability engineering, backend engineering, distributed systems, or operational tooling.
  • 3+ years of experience with architecture, design, system reliability, scalability, and technical delivery for production systems.
  • 2+ years of experience with open-source frameworks, modern engineering practices, or platform technologies.
  • 2+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments.
  • Demonstrated ownership of production systems operating in 24x7 environments.

Education
  • Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.

Additional Job Requirements
  • Ability to influence engineering outcomes within and across teams in complex organizations.
  • Must be able to communicate in a clear, concise, professional oral and written manner with customers, clients, co-workers, leadership, and other employees of the organization.
  • Must be able to perform effectively under pressure and in stressful situations, including during production support and high-severity incident response.
  • Must be able to participate in a 24x7 on-call rotation for incident response and production support of mission-critical platforms.


Annual Salary
$100,000.00 - $215,000.00
The above annual salary range is a general guideline. Multiple factors are taken into consideration to arrive at the final hourly rate/ annual salary to be offered to the selected candidate. Factors include, but are not limited to, the scope and responsibilities of the role, the selected candidate's work experience, education and training, the work location as well as market and business considerations.

At this time, GEICO will not sponsor a new applicant for employment authorization for this position.

About Geico

GEICO (Government Employees Insurance Company) is an American auto insurance company with headquarters in Chevy Chase, Maryland. It is the second largest auto insurer in the United States, after State Farm. GEICO is a wholly owned subsidiary of Berkshire Hathaway that provides coverage for more than 24 million motor vehicles owned by more than 15 million policy holders as of 2017. GEICO writes private passenger automobile insurance in all 50 U.S. states and the District of Columbia. The insurance agency sells policies through local agents, called GEICO Field Representatives, and over the phone directly to the consumer, and through their website.
Learn more about Geico
Size
40,000 employees
Industry
Founded
1936

Similar Jobs

More Jobs at Geico

More Enterprise Technology Jobs

Find similar Senior Engineer jobs: