Prove

Site Reliability Engineer

Prove$120K — $185K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years in Site Reliability or Platform Engineering
  • Deep expertise with AWS cloud platforms
  • Strong proficiency in Kubernetes and container orchestration
  • Experience with Terraform and infrastructure-as-code tools
  • Bachelor's degree in Computer Science, Engineering, or similar experience

Responsibilities

  • Optimize systems for reliability, performance, and scalability
  • Automate operations to improve efficiency and reduce toil
  • Ensure security compliance with least-privilege access controls
  • Implement cost-effective, efficient infrastructure
  • Participate in a 24/7 on-call rotation
  • Conduct post-incident reviews and implement preventive measures
  • Utilize observability data for system improvements

Benefits

  • Competitive salaries & Bonus Plan and Equity Plan
  • Modern Health for wellness support
  • 401(k) Retirement Plan & Match and Local Country Pension
  • Unlimited Vacation and Flexible hours
  • Comprehensive medical benefits for employees and families
  • Access to wellness services
  • Daily GrubHub stipend for lunch when in office
  • Positive and collaborative work environment
Full Job Description
Title: Site Reliability Engineer, Senior Site Reliability Engineer

Department: Platform Engineering

Reports To: Director, Platform Engineering

FLSA Status: Exempt

Location: US Remote

Job Summary

We are seeking Mid to Senior level Site Reliability Engineers to join our Platform Engineering team. In this role, you will be instrumental in designing, implementing, maintaining and deploying highly available complex, scalable and reliable systems leveraging automation, effective monitoring and infrastructure-as code. Working closely with our application engineering teams to ensure our services meet the highest standards of reliability, performance, and security.

Key Responsibilities for Senior level

The Site Reliability Engineering teams at Prove are responsible for driving maximum uptime for existing and developing products. Qualified candidates will be well versed in the difference between methods and ownership of outcomes and be able to demonstrate and document their relevant experience.

Observability Leadership
  • Design and implement comprehensive observability solutions across our infrastructure and within applications
  • Establish metrics, logging, and tracing systems that enable quick identification and resolution of issues
  • Create alerting thresholds and automated responses based on service level objectives (SLOs)
  • Provide actionable insights into service to service communications


Infrastructure Management
  • Design, build, and maintain scalable cloud infrastructure on AWS
  • Implement infrastructure-as-code using tools such as Terraform
  • Automate routine operational tasks to reduce toil and improve efficiency
  • Ensure infrastructure security compliance and implement least-privilege access controls
  • Design and implement infrastructure-as-code deployments for container based applications
  • Scale containers based on custom metrics for applications and critical observability infrastructure

Incident Response
  • Conduct thorough post-incident reviews and implement preventative measures
  • Use observability data to perform root cause analysis and system improvements
  • Participate in a 24/7 on call rotation to achieve 99.999% system availability.

Required Qualifications for Senior level
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering or equivalent experience. Software Engineering roles with a strong infrastructure and production engineering aspect also qualify.
  • Expert knowledge of observability platforms and practices (OpenTelemetry, Prometheus, Grafana, Jaeger, ELK stack / Splunk, etc)
  • Experience with Kubernetes and container orchestration
  • Strong experience with infrastructure-as-code tools (Terraform, Spacelift, Pulumi)
  • Proficiency in at least one programming language ( Go, Python )
  • Deep understanding of cloud platforms, preferably AWS
  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience


Key Responsibilities for IC3

The Site Reliability Engineering teams at Prove are responsible for driving maximum uptime for existing and developing products. Qualified candidates will be well versed in the difference between methods and ownership of outcomes and be able to demonstrate and document their relevant experience.

Preferred Qualifications
  • Experience with distributed systems and microservice architectures
  • Experience working in a high compliance environment
  • Hand-on experience instrumenting code with OpenTelemetry
  • Familiarity with service mesh technologies
  • Contributions to open-source projects
  • Experience in the identity verification or financial technology industry
  • Application development experience

Optimize
  • Improve new and existing systems by increasing reliability, performance, and scalability
  • Automate routine operational tasks to reduce toil and improve efficiency
  • Ensure infrastructure security compliance and implement least-privilege access controls
  • Implement efficient infrastructure that balances rapid development and cost
  • Embrace technological changes and development practices while maintaining reliability

Respond
  • Participate in a 24/7 on-call rotation
  • Conduct thorough post-incident reviews and implement preventative measures
  • Use observability data to identify system improvements

Run
  • Implement infrastructure as code in a myriad of high compliance development, production, and other environments
  • Scale developer experiences by being the standard bearer of an opinionated platform approach


Required Qualifications for IC3
  • 3+ years of experience in Site Reliability or Platform Engineering teams
  • Deep understanding of cloud platforms, particularly AWS
  • Strong experience with Kubernetes and container orchestration
  • Experience withTerraform and infrastructure-as-code tools
  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience


Preferred Qualifications
  • Experience with distributed systems and microservice architectures
  • Experience working in a high compliance environment
  • Experience with holistic monitoring and alerting for developing platforms
  • Skilled proficiency in at least one programming language (Go, Python)

Benefits & Perks for FTE Provers:
  • Competitive salaries & Bonus Plan (for eligible roles) and Equity Plan
  • Modern Health for financial, mental, and physical wellness
  • 401(k) Retirement Plan & Match (US Offices) and Local Country Pension (International Offices)
  • Unlimited Vacation and Flexible hours
  • Comprehensive medical benefits for you and your family ♥
  • Emotional & Physical Wellness - Access to wellness services (EAP & Prove Well-Being Reimbursement)
  • Bottomless snacks & beverages for certain office locations
  • Daily GrubHub stipend for lunch if coming into the office (US Offices)
  • A great place to work and connect with other talented Provers like yourself!


This position description should not be considered the final description of the position. The position description is not intended to be an all-inclusive list of duties and standards of the positions. It should be assumed that we would, to some extent, structure responsibilities in accordance with the successful candidate's capabilities and changing business conditions. Incumbents will follow any other instructions, and perform any other related duties, as assigned by their supervisor.

Site Reliability Engineer:
Metro 2: $130,000 - 150,000
Metro 3: $120,000 - 135,000

Senior Site Reliability Engineer:
Metro 2: $166,000 - 185,000
Metro 3: $153,000 - 171,000

Plus variable commission / company bonus. Offered salary will be determined by the applicant's education, experience, knowledge, skills, geo-location and abilities, as well as internal equity and alignment with market data.

Prove follows a market driven compensation philosophy based on geographic location and respective market rates. Job offers will be aligned to location. Please speak with your recruiter if you have questions. Prove defines:
  • Metro 2 - NYC metro area, Seattle metro area, Los Angeles metro area, and the Miami metro area.
  • Metro 3 - all other cities across the domestic United States, with the exception of the San Francisco Bay Area.

Benefits & Perks for FTE Provers:
  • Competitive salaries & Bonus Plan (for eligible roles) and Equity Plan
  • Modern Health for financial, mental, and physical wellness
  • 401(k) Retirement Plan & Match (US Offices) and Local Country Pension (International Offices)
  • Unlimited Vacation and Flexible hours
  • Comprehensive medical benefits for you and your family ♥
  • Emotional & Physical Wellness - Access to wellness services (EAP & Prove Well-Being Reimbursement)
  • Bottomless snacks & beverages for certain office locations
  • Daily GrubHub stipend for lunch if coming into the office (US Offices)
  • A great place to work and connect with other talented Provers like yourself!

About Prove

Wikipedia wik-ih-PEE-dee-? or /?w?ki-/ wik-ee-) is a free content, multilingual online encyclopedia written and maintained by a community of volunteers through a model of open collaboration, using a wiki-based editing system. Individual contributors, also called editors, are known as Wikipedians. It is the largest and most-read reference work in history. It is consistently one of the 15 most popular websites ranked by Alexa; as of 2021, Wikipedia was ranked the 13th most popular site. It is hosted by the Wikimedia Foundation, an American non-profit organization funded mainly through donations. On January 15, 2001, Jimmy Wales and Larry Sanger launched Wikipedia; Sanger coined its name as a portmanteau of "wiki" and "encyclopedia." Wales was influenced by the "spontaneous order" ideas associated with Friedrich Hayek and the Austrian School of economics, after being exposed to these ideas by Austrian economist and Mises Institute Senior Fellow Mark Thornton. Initially available only in English, versions in other languages were quickly developed. Its combined editions comprise more than 58 million articles, attracting around 2 billion unique device visits per month and more than 17 million edits per month as of November 2020. In 2006, Time magazine stated that the policy of allowing anyone to edit had made Wikipedia the "biggest encyclopedia in the world."
Learn more about Prove
Industry
Founded
2008

Similar Jobs

More Jobs at Prove

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: