Role Overview:We are seeking a Platform Engineer to contribute to the design, security, and scaling of our infrastructure on Google Cloud Platform (GCP). This role involves working across Kubernetes, infrastructure-as-code, security, compliance, and data platform tooling, focusing on maintaining system reliability, observability, and ease of use for product teams. It combines Site Reliability Engineering (SRE) practices with platform engineering, emphasizing both reliability improvements and infrastructure development.
Key Responsibilities:- Design, build, and maintain infrastructure on GCP using Terraform, ensuring environments are reproducible, version-controlled, and easy to audit.
- Operate and improve GKE clusters, including workload scheduling, autoscaling, networking, and container security hardening.
- Apply SRE principles to production systems: define and track SLOs/SLIs, build dashboards and alerting, participate in on-call rotation, and lead or contribute to incident response and postmortems.
- Implement and maintain IAM policies, service account permissions, and security controls aligned with least-privilege principles and compliance requirements (e.g., SOC 2, HIPAA, or similar depending on industry).
- Support and extend our data platform infrastructure, including BigQuery datasets, Dataflow pipelines, and related access controls and cost governance.
- Improve observability across the platform using tools like Cloud Monitoring, Cloud Logging, Prometheus, and Grafana.
- Partner with software engineering teams to provide self-service infrastructure patterns, reusable Terraform modules, and clear platform documentation.
- Participate in capacity planning, cost optimization, and disaster recovery planning for GCP workloads.
Required Skills:- 3-5 years of experience in a GCP platform engineering, DevOps, SRE, or infrastructure engineering role, with meaningful hands-on experience in GCP specifically (or strong equivalent experience in another major cloud with GCP exposure).
- Solid working knowledge of Kubernetes, including deployments, networking (VPC-native clusters, ingress, network policies), and container security practices.
- Experience writing and maintaining production Terraform code, including module design, state management, and multi-environment workflows.
- Practical understanding of IAM design on GCP: custom roles, service accounts, workload identity, organization policies, and audit logging.
- Familiarity with BigQuery and Dataflow (or similar data pipeline tooling), including access controls, cost management, and basic pipeline troubleshooting.
- Experience with SRE practices: SLO/SLI definition, alerting design, on-call participation, and incident postmortems.
- Comfort with scripting (Python, Bash, or Go) for automation and tooling work.
- Strong communication skills and the ability to document systems clearly for both technical and non-technical stakeholders.
Qualifications:- 4-6+ years experience required.
Preferred Skills:- GCP Professional certifications (Cloud Architect, DevOps Engineer, or Security Engineer).
- Experience with CI/CD tooling such as Cloud Build, GitHub Actions, or GitLab CI in a GCP context.