About the Role & TeamThe SRE team at PENN Entertainment is looking for a Senior Site Reliability Engineer to help build and operate the infrastructure behind a large-scale sports betting and media platform. You'll own critical infrastructure across compute, networking, storage, and cloud services (GCP/AWS) - driving complex migrations, building platform tooling and automation (ArgoCD, Helm, GitHub Actions), and improving observability and incident response across hundreds of production services spanning multiple regulated jurisdictions. We're looking for someone with strong Kubernetes and distributed systems experience, proficiency in Go, Python, or Bash, and a track record of leading cross-team infrastructure projects with high autonomy. You'll solve ambiguous problems, reduce operational toil through automation, mentor teammates, and bring a production-first perspective to architecture decisions - all on a team that values pragmatic engineering, real ownership, and continuous improvement of how we work.
About the Work- Drive complex infrastructure migrations and projects - scoping, planning, execution, and validation across multiple production environments and jurisdictions
- Build and maintain platform tooling and automation - ArgoCD, Helm, GitHub Actions, release pipelines, and service onboarding workflows that reduce toil for SRE and development teams
- Support development teams - consult on infrastructure needs, unblock cross-team dependencies, review architecture proposals, and help teams adopt platform tooling and best practices
- Design and improve observability and alerting - Datadog monitors, dashboards, and runbooks that surface meaningful signals and make systems operable by the whole team
- Provide operational support and incident response - investigate and resolve production issues through structured debugging and root cause analysis, and contribute to on-call rotations to maintain platform reliability
About You- 5+ Years of Experience in a similar role (DevOps, Site Relatability Engineer)
- Strong experience operating and troubleshooting Kubernetes in a production Linux environment (cluster lifecycle, networking, storage, scheduling)
- Experience working with AWS, GCP, and/or on-premise environments
- Proficiency in at least two of: Go, Python, Bash/Shell - for building tooling, automation, and debugging production systems
- Deep understanding of distributed systems - failure modes, networking fundamentals, capacity planning, and performance analysis
- Experience with GitOps and CI/CD workflows (ArgoCD, Helm, GitHub Actions, or similar)
- Experience with infrastructure-as-code (Terraform, Helm, or equivalent)
- Track record of leading complex migrations or infrastructure projects with cross-team dependencies
- Strong incident response and troubleshooting skills - structured debugging across multiple services and infrastructure layers
- Clear technical communication - you write docs others can use, explain trade-offs to varied audiences, and proactively unblock cross-team dependencies
Nice to have- Experience with service mesh technologies (Istio, Cilium)
- Familiarity with distributed storage systems (Ceph, or similar)
- Experience with bare-metal Kubernetes or Talos OS
- Exposure to regulated environments (sports betting, fintech, or similar compliance-heavy domains)
- Experience with Datadog or comparable observability platforms at scale
- Familiarity with database operations - PostgreSQL, connection pooling (PgBouncer), or database migration tooling
What We Offer- Competitive compensation package.
- Comprehensive Benefits package.
- Fun, relaxed work environment.
- Education and conference reimbursements.
#LI-REMOTE
Salary Range
$145,000-$193,000 CAD