OVERVIEWLevel99 is looking for a Sr. DevOps Engineer to build and operate the platform behind a growing, multi-location entertainment business. You'll run our cloud environment - CI/CD pipelines, Infrastructure-as-Code, observability, and security - while also managing the fleet of systems deployed across our venues using Ansible and AWX. It's a role that spans both sides: the infrastructure behind our cloud applications and the systems running on the floor, where things have to stay up while guests are playing games. You'll also help shape the standards and practices for how we build and operate as we continue to grow.
RESPONSIBILITIES- Ship reliably - build and maintain CI/CD pipelines and Infrastructure-as-Code that make deployments repeatable and safe
- Build and run our cloud platform - design and maintain the infrastructure behind our applications, from deployment architecture to networking
- Run the venue fleet - manage the systems deployed across our locations with Ansible and AWX, keeping configuration and updates consistent as we open new venues
- Support what runs on the floor - manage application dependencies and drivers on venue systems, partnering with our infrastructure and hardware teams on new deployments
- Keep systems healthy - instrument monitoring, logging, and alerting so problems get caught before guests notice
- Protect the environment - manage IAM, secrets, and access controls, and stay ahead of vulnerabilities
- Set the bar - establish documentation, runbooks, and standards that make the platform easier to operate as we grow
MUST-HAVE SKILLS- Hands-on production experience with a major cloud provider (GCP preferred; strong AWS or Azure experience considered)
- Production experience with containers and orchestration (Docker, Kubernetes)
- Production experience with Ansible (or similar) for configuration management across distributed systems
- Proficiency with Infrastructure-as-Code, Terraform preferred
- Experience building and maintaining CI/CD pipelines (CircleCI, GitLab CI, Jenkins, or similar)
- Scripting and automation ability in Python or Bash
- Comfort working with the Linux command line, including package management and troubleshooting application dependencies
- Demonstrated ownership of production systems, including on-call and incident response
- Strong written communication and documentation habits
OTHER DESIRABLE (BUT NOT NECESSARY) SKILLS & EXPERIENCE INCLUDE- Experience with monitoring and logging tooling (Datadog, Grafana/Prometheus, or similar)
- Experience with AWX or Ansible Automation Platform for fleet management at scale
- Working knowledge of cloud networking and IAM/secrets management
- Experience supporting systems that interface with physical hardware - device drivers, peripherals, or display and audio subsystems
- Database operations experience - backup strategy, restore testing, version upgrades
- Experience with distributed or multi-site environments, including compute running outside a central cloud region
$150,000 - $170,000 a year
While we don't expect a candidate to have deep experience in all of the above, we're looking for someone with the passion and capability to learn quickly in the areas that are new!YOU MIGHT BE A FIT ON THE LEVEL99 TEAM IF YOU...- Like to laugh, would be described as a "low maintenance, low drama" person, have a tendency to have a bit of fun while you work
- Have a high tolerance for ambiguity, like to go fast, and are excited to learn on the job
- Are just a little bit obsessive about getting the details right the first time
- Have a high energy personality, the kind of person who is typically smiling, and likes to "get it done now"