HOW YOU WILL IMPACT SKYLOAs a Senior, Cloud Infrastructure and Networking, in the Global Product Support & Customer Success organization, you are the Cloud Infrastructure domain authority within Skylo's production NTN network. Everything runs on the infrastructure you keep healthy - RAN NFs, Core NFs, OSS, BSS, and the observability pipeline itself. When a GKE node fails, when ArgoCD drifts, when a Persistent Volume Claim goes unavailable, when a PostgreSQL replica falls behind, when Prometheus WAL corrupts - you own the response.
You operate across Skylo's full hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). You own 24x7 platform health, the observability pipeline (Prometheus, VictoriaMetrics, Grafana, OpenTelemetry), persistent storage operations (PostgreSQL, Redis), and the operational interface with Network Implementation for all GitOps-driven infrastructure changes.
KEY RESPONSIBILITIESCloud Infrastructure Operations & Health Ownership- Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation across Skylo's GCP footprint.
- Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).
- Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation - distinguish transient platform events from systemic infrastructure degradation.
- Execute and own Cloud Infra runbooks for P2-P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation - without requiring engineering involvement for covered fault classes.
- Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks that reflect current cluster topology after every infrastructure change.
Observability Pipeline & Data Platform Operations- Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage and accuracy, OpenTelemetry collector health, and alert routing via Pub/Sub to the OSS.
- Maintain database reliability: PostgreSQL streaming replication health, backup and restore procedures, failover testing, query performance monitoring; Redis cluster operations, eviction policy management, and persistence configuration.
- Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, query performance, and completeness - the observability stack must be operational before the network events it monitors can be triaged.
- Partner with NI (Network Implementation & Infrastructure) on all planned infrastructure changes: receive advance notice, validate post-deployment observability, and sign off on operational readiness before the change window closes.
Cloud Infrastructure Incident Diagnosis & Escalation Authority- Serve as the L3 escalation authority for all Cloud Infra incidents: take ownership from the Incident Manager, diagnose at the Kubernetes, storage, network, and database layer using kubectl, GCP console, node logs, and infrastructure telemetry, and deliver a resolution or a decision-grade root cause.
- Lead Cloud Infra troubleshooting bridges: command the technical investigation for GKE node failures, cluster upgrade failures, storage outages, PubSub pipeline disruptions, database failover events, and ArgoCD sync failures - drive to resolution or clear engineering handoff.
- Diagnose and resolve infrastructure failure modes: node NotReady conditions, pod CrashLoopBackOff chains, PVC mount failures, CSI driver errors, network policy misconfigurations, Helm release drift, etcd latency spikes, and cross-cluster federation breaks.
- Participate in the global 24x7 on-call rotation as the Cloud Infra domain escalation tier - reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder.
SLO Engineering & Reliability- Define and maintain SLOs for all Cloud Infra components: GKE control plane availability, database query latency, storage IOPS, message pipeline throughput, and observability stack uptime - tied directly to network SLA commitments to MNO partners.
- Own error budget tracking and the process for trading error budget against deployment velocity; escalate when error budget burn rate requires engineering intervention or deployment freezes.
- Drive toil reduction: identify and eliminate manual Cloud Infra procedures; own the roadmap to automated cluster recovery, rolling restarts, storage repair, and certificate rotation in partnership with Ops Platform Engineering.
- Lead capacity planning for compute, storage, and network resources across public and private cloud - forecast growth based on subscriber projections and new MNO partner onboarding.
Root Cause Analysis & Post-Incident Ownership- Own Cloud Infra RCA end-to-end: lead the investigation, document the complete causal chain from infrastructure trigger through upstream NF impact, and deliver systemic action items with owners, timelines, and measurable success criteria.
- Deliver Initial RCA documentation within defined SLA windows; identify systemic infrastructure failure patterns - cluster upgrade regressions, storage controller bugs, network policy drift, resource exhaustion trends - and translate them into engineering requirements.
- Contribute to the weekly and monthly Network Performance Report: infrastructure availability, database latency trends, storage IOPS, observability pipeline health, and SLA deviation analysis.
Runbook Authorship & Operational Standards- Author, own, and maintain all Cloud Infra runbooks and SOPs: GKE node recovery, database failover, storage expansion, Prometheus WAL repair, ArgoCD rollback, certificate rotation, and cluster upgrade procedures - every procedure tested before production reliance.
- Define the diagnostic decision tree for each known infrastructure fault class: entry condition, triage steps, isolation method, resolution action, and escalation criteria - written at the level where a Senior NRE can execute independently.
- Validate and sign off on operational readiness for all infrastructure changes: DCI/DCE build-outs, GKE cluster expansions, Kubernetes version upgrades, and new on-premise hardware deployments.
Cross-Functional Collaboration & Team Development- Partner with OPE as the Cloud Infra domain's primary automation consumer: define Kubernetes event schemas, alert-to-action contracts, and closed-loop policy requirements for infrastructure auto-remediation.
- Represent Cloud Infra in NI architecture reviews: define observability and operational readiness requirements for all infrastructure expansions and GitOps pipeline changes before go-live.
- Collaborate with Core NRE and RAN NRE on infrastructure-layer issues affecting network functions: pod scheduling, PVC availability, network policy changes, and platform upgrade impacts on NF workloads.
- Partner with security teams on infrastructure hardening: patch compliance, RBAC policies, network segmentation, container image scanning, and runtime security monitoring across public and private cloud.
- Surface toil and automation opportunities to the Service Assurance & Automation team - document the procedure, frequency, and MTTR cost as structured input to the automation backlog.
- Mentor Senior NREs in Cloud Infra domain depth: Kubernetes troubleshooting patterns, storage operations, database reliability, observability pipeline internals, and escalation judgment.
GitOps, IaC & Platform Engineering Interface- Own operational oversight of GitOps tooling in production: ArgoCD sync health, Helm chart version management, drift detection, and rollback execution for multi-cluster deployments.
- Review and validate Infrastructure as Code (Terraform, Ansible) changes that impact production - ensure operational impact is assessed and observability is in place before merge.
- Engage Skylo's Platform Engineering and NI teams with full operational context when issues exceed operational resolution authority - deliver a structured problem statement, infrastructure telemetry bundle, and a clear question rather than a vague escalation.
REQUIRED QUALIFICATIONS- 5+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment - with direct on-call ownership for Kubernetes-at-scale environments.
- Deep Kubernetes expertise: multi-cluster operations (GKE or EKS), node pool management, RBAC, network policies, persistent storage (PVC, CSI drivers), CRD/operator patterns, and production cluster upgrade procedures.
- Hybrid cloud operations: hands-on experience operating both public cloud (GCP or AWS) and on-premise/private cloud infrastructure (bare-metal Kubernetes, KVM, or hyperconverged platforms).
- Production observability stack ownership: Prometheus (federation, remote write, WAL management), Grafana, VictoriaMetrics, OpenTelemetry, and alerting pipeline design with Pub/Sub or equivalent.
- Database reliability: PostgreSQL streaming replication, backup/restore, failover procedures, and performance tuning; Redis cluster operations and persistence management.
- GitOps tooling in production: ArgoCD or Flux CD for multi-cluster operations; Helm chart authorship and version management; Terraform or Ansible for infrastructure provisioning.
- SRE fundamentals: SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design.
- Container and Linux internals: container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding.
- Runbook authorship: ability to write infrastructure diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure.
- Strong written and verbal communication: capable of delivering RCA documents, engineering escalations with structured problem statements, and MNO-facing infrastructure summaries.
PREFERRED QUALIFICATIONS- Experience operating cloud infrastructure for telecom or NTN workloads: 5G Core NF hosting, vRAN compute requirements, or satellite ground segment infrastructure.
- Software-defined storage expertise: Ceph, Rook, or equivalent distributed storage systems at production scale.
- Private cloud platform experience: KubeVirt, Harvester, or OpenStack for VM-container convergence on bare-metal infrastructure.
- Networking depth: BGP routing, VXLAN overlays, EVPN fabrics, software-defined networking, and hardware load balancer operations.
- Strong development background in Go or Python for building custom automation tooling, Kubernetes operators, or infrastructure lifecycle integrations.
- FinOps experience: cloud cost optimization, resource lifecycle automation, and capacity right-sizing across public cloud footprints.
- Certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect.
WHAT WE OFFEROur worldwide culture encourages a flexible approach to work, and we offer an attractive range of benefits:
- Competitive compensation packages including a stock option-based equity program
- Comprehensive benefits including medical