About the RoleWe're hiring a Senior Software Engineer on the Infrastructure team to own the foundational infrastructure and internal developer platform that every other engineering team at Commure builds on. This is a horizontal team supporting multi-product infrastructure that you will design, build, and operate end-to-end:
- Cloud infrastructure and IaC
- Kubernetes fleet and workload orchestration
- Internal developer platform
- Release and deployment (GitOps)
- Service mesh, traffic management, and networking
- Observability: metrics, logs, and traces
- Zero-trust access and on-prem connectivity
The stack today runs on public cloud (GCP, AWS, and Azure), with all infrastructure defined as code (Terraform and controller-based). Argo CD drives deployment; Helm handles application packaging; Prometheus, Grafana, and OpenTelemetry power observability. Service mesh is an active build-out, and the shape of it is yours to define.
This is a hands-on IC role with broad scope. You'll make architectural calls, write the code that matters most, and set the patterns other teams build on.
What You'll DoYou'll own several of these verticals within the team's scope end-to-end.
- Build out the internal developer platform: golden-path templates, self-serve tooling, local development environments, and CI/CD pipelines.
- Own the cloud foundation: GCP, AWS, and/or Azure infrastructure managed as code with Terraform and controller-based provisioning via Kubernetes operators and Crossplane.
- Run the Kubernetes fleet: cluster lifecycle, upgrades, autoscaling, node management, and multi-cluster patterns. Shape how services are packaged and deployed with Helm.
- Design the traffic and network layer: service mesh, ingress, mTLS, and traffic management (routing, rate limiting, canary, circuit-breaking, RPC).
- Own the release and deployment story with Argo CD. GitOps workflows, progressive delivery (canary, blue-green), rollback safety, and environment promotion patterns.
- Own the observability stack: OpenTelemetry based instrumentation, metrics (Prometheus), dashboards (Grafana), distributed tracing, logging, unified alerting, templated dashboard, etc.
- Build out zero-trust access to internal and external systems: VPN, BeyondCorp, short-lived credentials, and on-prem connectivity.
- Partner with Security on secrets management, policy-as-code, and secure-by-default patterns that meet HIPAA and SOC 2 by default.
What You Have- 6+ years of software engineering experience in infrastructure, platform, or site reliability engineering roles.
- Experience building internal developer platforms with a strong product mindset (treating developers as customers).
- Experience with public cloud (GCP, AWS, Azure) and managed cloud services.
- Experience with on-prem or hybrid environments and data center / cloud migrations.
- Experience with cloud-native technologies.
- Experience with Infrastructure-as-Code (Terraform, Pulumi).
- Experience with plus controller-based infrastructure management (Crossplane, Kubernetes operators).
- Experience with service mesh technologies and software-defined networking (SDN).
- Experience with modern release workflows using GitOps and progressive delivery (Argo CD, Flux, Kargo).
- Experience with observability stack (Prometheus, Grafana, OpenTelemetry).
- Experience in regulated industries (healthcare, finance) with HIPAA and SOC 2 obligations.