Full Job Description
From a reliability standpoint, this role involves evaluating the scalability, resiliency, performance, and security properties and techniques used in production environments. It supports the uptime of production services through an On-Call rotation, which includes monitoring and alerting to meet internal Service Level Objectives (SLOs) and customer-facing Service Level Agreements (SLAs). Ensuring reliable incident processes is achieved by conducting Disaster Recovery drills. The role also focuses on improving reliability through incident management by investigating incidents, implementing remediation strategies, and learning from past incidents to make improvements. It involves determining the reliability and security requirements of components and systems to meet the reliability objectives of the company, customers, and any relevant governmental agencies. Additionally, the role aims to reduce operational expenses through automation, by identifying and mitigating failure points, and automating repetitive and resource-intensive tasks. It also involves developing new acceleration techniques and analytical tools to ensure the early identification of potential issues with new products, packaging, processes, and overall product reliability.
Ensures a reliable and scalable network, manages network / cloud infrastructure and storage systems supporting business operations, and responds to planned maintenance, real-time outages, and issues.
Plans, designs and implements local and wide-area network solutions between multiple platforms and protocols.
What You'll Do:
• 7+ yrs experienced professional using best practices and knowledge of internal or external business issues to improve products or services.
• Works independently, but receives minimal guidance and direction from leader then determines best approach to accomplish work.
• Acts as a resource for colleagues with less experience.
• Understands project and/or department needs and establishes relationships with appropriate cross-functional stakeholders to gather input, collect information, and complete work steps.
• Designs and deploys small to mid-size or moderately complex solutions to optimize reliability, availability, latency, and performance.
• Integrates knowledge of design, automation, and deployment with expertise in coding to improve service reliability for existing or new systems and adapts for regions, countries, or customers.
• Designs and tests high availability and disaster recovery measures for our services to ensure automation is improving reliability, scalability, and velocity.
• Forecasts and builds reports to determine at what point resources will be at capacity.
• Designs and implements tools that provide visibility into performance and reliability of our infrastructure.
• Builds automated platforms.
• Monitors the environment and works with Developers and Ops to identify problems and develop monitoring tools that provide visibility into performance and reliability, serves as on-call SRE, leads post mortems, and writes root cause analysis.
• Builds and ensures security controls are in place in regards to architectural design, collaborates with security in designing or providing input to security controls, and may actively contribute in security incident response.
Minimum Qualifications:
• Bachelors + 7 years of related experience, or Masters + 4 years of related experience, or PhD + 1 year of related experience.
• Requires solid conceptual and practical knowledge in primary technical job family and knowledge of related technical job families; has worked with a range of technologies.