Site Reliability Engineer

V2Soft

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in a relevant field.
  • 4+ years of IT experience with 3+ years in development.
  • Hands-on experience with Google Cloud Platform (GCP).
  • Proficiency in monitoring tools like Dynatrace or similar.
  • Familiarity with ITSM tools such as ServiceNow.

Responsibilities

  • Collaborate with Infrastructure teams to automate routine tasks.
  • Monitor and manage production environments, addressing issues proactively.
  • Build advanced tooling for system access monitoring and log session recording.
  • Work with engineering teams to enhance on-call efficiency and incident management.
  • Perform capacity planning to support growing demands.
  • Maintain and monitor systems for proactive health checks.
  • Continuously enhance system performance through data-driven optimization.

Benefits

  • Hybrid work schedule (Monday-Thursday in the office).
  • Opportunity to work on observability and monitoring of GCP data platforms.
  • Hands-on experience with cloud infrastructure and CI/CD pipelines.
Full Job Description
Position Description:

Only W2, No C2C.


Employees in this job function are responsible for ensuring availability, reliability and performance of cloud and network systems and services by automating routine manual tasks Key Responsibilities: 1. Collaborate with Infrastructure teams in implementing critical solutions by automating routine tasks 2. Monitor and manage production environments, proactively identifying and resolving issues. 3. Participate in building advanced tooling for system access monitoring, log session recording, administration of reliability across multiple geographically distributed data centers. 4. Engage with engineering teams to improve on-call efficiencies, drive incident management and post-mortem analysis.5. Perform capacity planning and optimization to support growing demands and traffic patterns. 6. Maintaining, monitoring and alerting systems for proactive system health checks. 7. Continuously improve system performance, stability, and security through data-driven analysis and optimization. 8. Facilitate knowledge sharing by creating and maintaining comprehensive documentation & diagrams
Skills Required:
Big Query, Dynatrace, GCP
Skills Preferred:
GCP Cloud Run, Python, Troubleshooting (Problem Solving)
Experience Required:
Engineer 2 Exp.: Practitioner: 1 coding language or framework. 4+ years in IT; 3+ years in development; Hands-on experience with Google Cloud Platform (GCP). Proficiency with monitoring/observability tools, ideally Dynatrace (or comparable, e.g., Datadog, New Relic). Familiarity with ITSM tools such as ServiceNow (incident, problem, change management)
Experience Preferred:
Familiarity with the use of AI tools - agents, skills, LLMs, copilot. Experience defining and tracking SLAs/SLOs/SLIs
Education Required:
Bachelor's Degree
Additional Information :
Hybrid - Monday -Thursday in the office We're looking for a Site Reliability Engineer to join our GDI&A SRE team, focused on observability, monitoring, and technical consulting across our GCP-based data platforms. You'll work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to keep critical systems healthy, performant, and reliable at scale.

Similar Jobs

More Jobs at V2Soft

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: