Fandom

Senior Cloud Engineer

Fandom$122K — $204K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Technical/Network Operations, DevOps, or SRE roles managing large-scale platforms (10M+ active users)
  • Deep proficiency in Linux administration, networking, and secure systems
  • Hands-on experience with Kubernetes, CI/CD pipelines, and monitoring systems
  • Experience managing production MySQL databases with replication and failover knowledge
  • Ability to use generative AI tools for productivity with skill in auditing outputs

Responsibilities

  • Design and scale high-availability routing architectures for global traffic
  • Automate cloud and edge infrastructure as code using Terraform and Chef
  • Maintain and optimize large-scale production environments
  • Develop tools for operational automation and security monitoring
  • Lead planning meetings and system evaluations for operational excellence
  • Participate in on-call rotation for incident response

Benefits

  • Vibrant team culture
  • Comprehensive Medical, Dental, Vision
  • Unlimited training resources (Udemy and more)
  • Flexible working hours and time off
  • Equity and retirement programs including 401K match
  • Paid parental leave
  • Startup culture in an international environment
Full Job Description
About this Role

We're looking for a Senior Cloud Engineer to help evolve and support the infrastructure that powers our platform for over 300 million fans around the world. This is a hands-on role focused on building reliable, scalable systems in a Linux and Kubernetes-based environment.

As part of the TechOps team, you'll report to the Manager of TechOps and work closely with developers, product engineers, and other infrastructure teams. You'll contribute to our CI/CD, monitoring, automation, and cloud efforts - helping ensure Fandom's platform remains fast, stable, and secure as we grow.

This is a great opportunity for someone who enjoys solving complex infrastructure challenges, improving deployment systems, and enabling engineering teams to move faster and safer.
You Will...
  • Design, architect, and scale high-availability routing architectures to seamlessly balance and secure global user traffic across a hybrid footprint of on-premise datacenters, AWS, and GCP.
  • Manage and automate cloud and edge infrastructure as code (IaC) using Terraform and Chef, ensuring consistent configurations for Kubernetes, global Cloudflare services, and CI/CD pipelines.
  • Maintain and optimize large-scale production environments, orchestrating high-availability MySQL replication, automated failover, robust disaster recovery, and system monitoring.
  • Develop internal tools for operational automation, system/data backups, performance tuning, and comprehensive security monitoring.
  • Drive operational excellence by leading planning meetings, retrospectives, and RCAs, while continuously evaluating systems against industry best practices.
  • Participate in an on-call rotation to maintain production stability, handle incident responses, and collaborate with/mentor cross-functional engineering teams.
You Have...
  • 5+ years of experience in Technical/Network Operations, DevOps, or SRE roles managing large-scale production platforms (e.g., 10M+ monthly active users).
  • Deep proficiency in Linux systems administration, networking protocols (TCP/IP, routing), secure systems practices, and scripting/programming (Go, Python, or Bash).
  • Proven hands-on experience with core infrastructure tech: Kubernetes/container orchestration, CI/CD pipelines (GitHub Actions, Jenkins), and monitoring/reliability systems (e.g., Prometheus).
  • Practical experience managing production MySQL database environments, including deep familiarity with replication topologies, failover mechanisms, and performance tuning.
  • Demonstrated capability using generative AI tools (e.g., Gemini, NotebookLM) to enhance productivity, paired with the ability to critically audit and verify outputs for accuracy, security, and context.
Bonus Points...
  • Advanced Cloudflare expertise, including CDN optimization, WAF security, DNS management, edge performance tuning, and Cloudflare Tunnels.
  • Strong understanding of distributed systems architecture, edge caching, and centralized log management using the ELK stack (Elasticsearch, Logstash, Kibana).
  • Experience defining SLOs and instrumentation, implementing meaningful metrics, logs, and traces to reduce alert noise and drive postmortem action items.
Benefits & Perks
  • Salary Range = $122k - $204k (Actual salary available will vary based on location and market factors.)
  • Vibrant team culture
  • Comprehensive Medical, Dental, Vision
  • Training (unlimited Udemy + more)
  • Flexible working hours and time off
  • Equity & Retirement Programs including 401K match
  • Paid Parental Leave
  • International work environment with start-up culture
#LI-TM1

About Fandom

Fandom is a wiki hosting service which hosts wikis mainly on entertainment. Its domain is operated by Fandom, Inc., a for-profit Delaware company founded in October 2004 by Jimmy Wales and Angela Beesley. Fandom was acquired in 2018 by TPG Capital and Jon Miller through Integrated Media Co. Fandom uses MediaWiki, the open-source wiki software used by Wikipedia. Fandom, Inc. derives its income from advertising and sold content, publishing most user-provided text under copyleft licenses. The company also runs the associated Fandom editorial project, offering pop-culture and gaming news. Fandom wikis are hosted under the domain fandom.com, but some, especially those that focus on subjects other than media franchises, were hosted under wikia.org until November 2021.
Learn more about Fandom
Industry
Founded
2004

Similar Jobs

More Jobs at Fandom

More Information Technology Jobs

Find similar Senior Cloud Engineer jobs: