To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.
Job Category
Software Engineering
Job Details
Slack is looking for an Engineering Manager on the Core Services engineering team. The purpose of this role is to keep Slack working seamlessly at massive scale — it's the infrastructure that keeps Slack feeling lightning fast. This person will be a steward of some of the most critical, high-throughput systems at Slack — the kind of work where doing it well means millions of users, at the world’s most important and interesting companies, never experience a hiccup.
About the OpportunityWe operate at tremendous scale with systems that process millions of events per second. Our team maintains and builds performant services that include:
- Edge Caches
- Realtime Services
- Enterprise Key Management (EKM)
- International Data Residency (IDR)
- Image Uploads, Downloads, and Proxying
We've done our job correctly when none of our users think about us, much like a vital utility. We don't typically ship new user-facing features, but rather ensure our systems are exceptionally performant, highly available, reliable, and scalable. In other words, we make Slack work seamlessly. Slack's backend services are written in Go and Java, and Slack's API is built on PHP/Hack.
The Core Services TeamWe are a team building critical services that enable Slack to scale for our largest customers. Our edge caches power the most latency-critical queries for our largest entities today, serving approximately 35% of all Slack API traffic. Our Realtime Services (RTS) power Slack's core messaging stack — running 10 services, 4 of which are mission-critical Tier 0 services — ensuring that messages, presence, and real-time events are delivered reliably at scale. We also own services that power Slack's Enterprise Key Management and International Data Residency features, enabling Slack to win and retain enterprise and government customers with strict compliance requirements (FedRAMP/NIST).
We have a strong commitment to quality and understand that simplicity and reliability are the foundations of our work.
We are actively building the future of edge services at Slack, so we can support a growing number of globally distributed customers with their compliance requirements with minimal latency.
What You'll Be Doing:- Owning the technical direction and roadmap across the Core Services portfolio: Edge Caches, Realtime Services, Enterprise Key Management, International Data Residency, and Image Proxying and Uploads. You’ll have large footprint and the opportunity to drive major impact. And you'll decide what to invest in now, what to sequence later, and what to decline — balancing reliability, migration, and new capability work.
- Leading a multi-system migration of Flannel to Loom, our next-generation edge cache, while keeping the current systems running at full reliability for millions of users. You'll drive the team to ship new capabilities without introducing risk to infrastructure that people depend on daily.
- Driving delivery and execution while maintaining reliability as a top priority, including sprint planning, product launches, incident follow-ups, operational sustainability analyses, and building the technical case for investment with leadership. You'll ensure the team doesn't trade reliability for velocity, but is always high-performing.
- Owning on-call operations and incident management for Tier 0 services. You'll manage on-call rotations across partnering teams, triage production issues, lead incident response for your services, and drive improvements to runbooks, alerting, and operational tooling.
- Partnering cross-functionally with Product, Sales, Compliance, Security, AWS, and other infrastructure teams on dependencies, timelines, and risks. You'll be the shortest path to solutions — whether it's unblocking a compliance audit for South Korea, coordinating IDR region launches with Sales and AWS, or negotiating priorities with partner teams for shared work.
- Representing your team's work in leadership reviews, quarterly and annual planning, and executive forums. You'll build compelling business-and-technical cases for your team's investments and tell your leadership what's really going on early enough that nothing becomes a fire drill.
- Building and maintaining a high-trust team culture where engineers do impactful work and grow. You'll hire well, manage performance with fairness and transparency, mentor emerging technical leaders, and celebrate your team's wins.
- Managing cost and capacity for services that run at massive scale, working with finance on AWS spend trends, exec-approved budget increases, and cost-saving initiatives like the Loom migration (projected $1-2M savings).
- Driving the team to adopt AI development tooling for maximum impact, while maintaining engineering rigor and quality for critical infrastructure.
You're Our Person If: - You've led engineering teams for 5+ years building and operating high-scale, high-availability backend infrastructure or distributed systems.
- You have strong software engineering fundamentals with previous experience as a Senior+ backend engineer, with enough depth to be credible in architecture discussions about caching, distributed systems, observability and reliability best practices, and system design — and to build SME knowledge in a new area fast.
- You have a history of managing Tier 0 / Tier 1 services where uptime matters deeply, including owning on-call rotations, leading incident response, and driving operational improvements.
- You have experience with realtime messaging systems or event-driven architectures at scale — presence, message delivery, pub/sub, or similar systems where latency and reliability directly impact user experience.
- You have experience leading multi-quarter migration or platform evolution projects while simultaneously keeping existing systems reliable.
- You've managed teams through prioritization trade-offs — weighing reliability vs. feature work vs. cost vs. compliance — and can articulate those decisions to both engineers and executives.
- You have experience partnering with cross-functional stakeholders (Product, Sales, Compliance, Security, Infrastructure) to turn business requirements into engineering plans.
- You understand how to build a high-performing team culture: hiring, performance management, retention of top performers, and growing new technical leaders.
- You have experience managing engineers across multiple time zones, with an async rhythm that doesn't depend on overlap hours.
- An understanding of how and where AI software development tooling can be deployed for maximum team impact.
Even Better If: - You have experience with the Go and/or Java programming language at scale.
- Experience with AWS infrastructure, including multi-region architectures.
- Experience with Kubernetes and container-based deployments.
- Experience working within highly regulated environments where an understanding of FedRAMP/NIST frameworks was essential.
- Experience with Enterprise Key Management, data residency, or compliance-driven infrastructure.
- You've managed teams responsible for Edge Caching, CDN, or API gateway systems.
- You've led teams through system deprecation and migration while maintaining service reliability.