Senior Engineer, Cloud Infrastructure and Networking

Skylo Technologies

• $125K — $135K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in infrastructure engineering, Site Reliability Engineering, or cloud operations in a 24x7 production environment.
  • Deep Kubernetes expertise including multi-cluster operations and persistent storage management.
  • Hands-on experience with hybrid cloud operations in both public (GCP, AWS) and on-premise environments.
  • Proven ownership of production observability stacks like Prometheus and Grafana.
  • Strong database reliability skills with PostgreSQL and Redis operations.
  • Experience with GitOps tooling in production environments, particularly ArgoCD and Terraform.
  • Solid understanding of SRE principles including SLO/SLI/SLA definitions and error budget management.
  • Strong communication skills for effective documentation and engineering interfaces.

Responsibilities

  • Own 24x7 health of Skylo's hybrid cloud infrastructure including GKE and on-premise Kubernetes clusters.
  • Monitor and triage infrastructure alarms, distinguishing transient events from systemic risks.
  • Execute and document runbooks for common infrastructure faults and failures without needing engineering input.
  • Drive root cause analysis for incidents, producing documentation and actionable items to prevent recurrence.
  • Coordinate cross-functional efforts and lead troubleshooting bridges during significant incidents.
  • Define and maintain service level objectives for cloud infrastructure components.
  • Collaborate with automation teams to identify toil reduction opportunities and standardize procedures.

Benefits

  • Flexible work culture promoting a balanced approach to work.
  • Comprehensive medical benefits coverage.
  • Stock option-based equity program for employees.
Full Job Description
HOW YOU WILL IMPACT SKYLO

As a Senior, Cloud Infrastructure and Networking, in the Global Product Support & Customer Success organization, you are the Cloud Infrastructure domain authority within Skylo's production NTN network. Everything runs on the infrastructure you keep healthy - RAN NFs, Core NFs, OSS, BSS, and the observability pipeline itself. When a GKE node fails, when ArgoCD drifts, when a Persistent Volume Claim goes unavailable, when a PostgreSQL replica falls behind, when Prometheus WAL corrupts - you own the response.

You operate across Skylo's full hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). You own 24x7 platform health, the observability pipeline (Prometheus, VictoriaMetrics, Grafana, OpenTelemetry), persistent storage operations (PostgreSQL, Redis), and the operational interface with Network Implementation for all GitOps-driven infrastructure changes.

KEY RESPONSIBILITIES

Cloud Infrastructure Operations & Health Ownership
  • Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation across Skylo's GCP footprint.
  • Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).
  • Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation - distinguish transient platform events from systemic infrastructure degradation.
  • Execute and own Cloud Infra runbooks for P2-P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation - without requiring engineering involvement for covered fault classes.
  • Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks that reflect current cluster topology after every infrastructure change.


Observability Pipeline & Data Platform Operations
  • Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage and accuracy, OpenTelemetry collector health, and alert routing via Pub/Sub to the OSS.
  • Maintain database reliability: PostgreSQL streaming replication health, backup and restore procedures, failover testing, query performance monitoring; Redis cluster operations, eviction policy management, and persistence configuration.
  • Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, query performance, and completeness - the observability stack must be operational before the network events it monitors can be triaged.
  • Partner with NI (Network Implementation & Infrastructure) on all planned infrastructure changes: receive advance notice, validate post-deployment observability, and sign off on operational readiness before the change window closes.


Cloud Infrastructure Incident Diagnosis & Escalation Authority
  • Serve as the L3 escalation authority for all Cloud Infra incidents: take ownership from the Incident Manager, diagnose at the Kubernetes, storage, network, and database layer using kubectl, GCP console, node logs, and infrastructure telemetry, and deliver a resolution or a decision-grade root cause.
  • Lead Cloud Infra troubleshooting bridges: command the technical investigation for GKE node failures, cluster upgrade failures, storage outages, PubSub pipeline disruptions, database failover events, and ArgoCD sync failures - drive to resolution or clear engineering handoff.
  • Diagnose and resolve infrastructure failure modes: node NotReady conditions, pod CrashLoopBackOff chains, PVC mount failures, CSI driver errors, network policy misconfigurations, Helm release drift, etcd latency spikes, and cross-cluster federation breaks.
  • Participate in the global 24x7 on-call rotation as the Cloud Infra domain escalation tier - reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder.


SLO Engineering & Reliability
  • Define and maintain SLOs for all Cloud Infra components: GKE control plane availability, database query latency, storage IOPS, message pipeline throughput, and observability stack uptime - tied directly to network SLA commitments to MNO partners.
  • Own error budget tracking and the process for trading error budget against deployment velocity; escalate when error budget burn rate requires engineering intervention or deployment freezes.
  • Drive toil reduction: identify and eliminate manual Cloud Infra procedures; own the roadmap to automated cluster recovery, rolling restarts, storage repair, and certificate rotation in partnership with Ops Platform Engineering.
  • Lead capacity planning for compute, storage, and network resources across public and private cloud - forecast growth based on subscriber projections and new MNO partner onboarding.


Root Cause Analysis & Post-Incident Ownership
  • Own Cloud Infra RCA end-to-end: lead the investigation, document the complete causal chain from infrastructure trigger through upstream NF impact, and deliver systemic action items with owners, timelines, and measurable success criteria.
  • Deliver Initial RCA documentation within defined SLA windows; identify systemic infrastructure failure patterns - cluster upgrade regressions, storage controller bugs, network policy drift, resource exhaustion trends - and translate them into engineering requirements.
  • Contribute to the weekly and monthly Network Performance Report: infrastructure availability, database latency trends, storage IOPS, observability pipeline health, and SLA deviation analysis.


Runbook Authorship & Operational Standards
  • Author, own, and maintain all Cloud Infra runbooks and SOPs: GKE node recovery, database failover, storage expansion, Prometheus WAL repair, ArgoCD rollback, certificate rotation, and cluster upgrade procedures - every procedure tested before production reliance.
  • Define the diagnostic decision tree for each known infrastructure fault class: entry condition, triage steps, isolation method, resolution action, and escalation criteria - written at the level where a Senior NRE can execute independently.
  • Validate and sign off on operational readiness for all infrastructure changes: DCI/DCE build-outs, GKE cluster expansions, Kubernetes version upgrades, and new on-premise hardware deployments.


Cross-Functional Collaboration & Team Development
  • Partner with OPE as the Cloud Infra domain's primary automation consumer: define Kubernetes event schemas, alert-to-action contracts, and closed-loop policy requirements for infrastructure auto-remediation.
  • Represent Cloud Infra in NI architecture reviews: define observability and operational readiness requirements for all infrastructure expansions and GitOps pipeline changes before go-live.
  • Collaborate with Core NRE and RAN NRE on infrastructure-layer issues affecting network functions: pod scheduling, PVC availability, network policy changes, and platform upgrade impacts on NF workloads.
  • Partner with security teams on infrastructure hardening: patch compliance, RBAC policies, network segmentation, container image scanning, and runtime security monitoring across public and private cloud.
  • Surface toil and automation opportunities to the Service Assurance & Automation team - document the procedure, frequency, and MTTR cost as structured input to the automation backlog.
  • Mentor Senior NREs in Cloud Infra domain depth: Kubernetes troubleshooting patterns, storage operations, database reliability, observability pipeline internals, and escalation judgment.


GitOps, IaC & Platform Engineering Interface
  • Own operational oversight of GitOps tooling in production: ArgoCD sync health, Helm chart version management, drift detection, and rollback execution for multi-cluster deployments.
  • Review and validate Infrastructure as Code (Terraform, Ansible) changes that impact production - ensure operational impact is assessed and observability is in place before merge.
  • Engage Skylo's Platform Engineering and NI teams with full operational context when issues exceed operational resolution authority - deliver a structured problem statement, infrastructure telemetry bundle, and a clear question rather than a vague escalation.


REQUIRED QUALIFICATIONS
  • 5+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment - with direct on-call ownership for Kubernetes-at-scale environments.
  • Deep Kubernetes expertise: multi-cluster operations (GKE or EKS), node pool management, RBAC, network policies, persistent storage (PVC, CSI drivers), CRD/operator patterns, and production cluster upgrade procedures.
  • Hybrid cloud operations: hands-on experience operating both public cloud (GCP or AWS) and on-premise/private cloud infrastructure (bare-metal Kubernetes, KVM, or hyperconverged platforms).
  • Production observability stack ownership: Prometheus (federation, remote write, WAL management), Grafana, VictoriaMetrics, OpenTelemetry, and alerting pipeline design with Pub/Sub or equivalent.
  • Database reliability: PostgreSQL streaming replication, backup/restore, failover procedures, and performance tuning; Redis cluster operations and persistence management.
  • GitOps tooling in production: ArgoCD or Flux CD for multi-cluster operations; Helm chart authorship and version management; Terraform or Ansible for infrastructure provisioning.
  • SRE fundamentals: SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design.
  • Container and Linux internals: container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding.
  • Runbook authorship: ability to write infrastructure diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure.
  • Strong written and verbal communication: capable of delivering RCA documents, engineering escalations with structured problem statements, and MNO-facing infrastructure summaries.


PREFERRED QUALIFICATIONS
  • Experience operating cloud infrastructure for telecom or NTN workloads: 5G Core NF hosting, vRAN compute requirements, or satellite ground segment infrastructure.
  • Software-defined storage expertise: Ceph, Rook, or equivalent distributed storage systems at production scale.
  • Private cloud platform experience: KubeVirt, Harvester, or OpenStack for VM-container convergence on bare-metal infrastructure.
  • Networking depth: BGP routing, VXLAN overlays, EVPN fabrics, software-defined networking, and hardware load balancer operations.
  • Strong development background in Go or Python for building custom automation tooling, Kubernetes operators, or infrastructure lifecycle integrations.
  • FinOps experience: cloud cost optimization, resource lifecycle automation, and capacity right-sizing across public cloud footprints.
  • Certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect.


WHAT WE OFFER

Our worldwide culture encourages a flexible approach to work, and we offer an attractive range of benefits:
  • Competitive compensation packages including a stock option-based equity program
  • Comprehensive benefits including medical

Similar Jobs

More Jobs at Skylo Technologies

More Information Technology Jobs

Find similar Senior Engineer, Cloud Infrastructure and Networking jobs: