Radar.io

Engineering Manager, Site Reliability

Radar.io$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in software engineering and site reliability engineering management
  • Strong background in AWS environment and Terraform usage
  • Knowledge of Kubernetes and managing multi-region production infrastructure
  • Experience with MongoDB in high-availability, large-scale applications
  • Familiarity with observability and incident management tools

Responsibilities

  • Manage and mentor a distributed SRE team to enhance infrastructure reliability
  • Oversee observability, monitoring, and incident management practices
  • Utilize Terraform for core infrastructure management and deploy on AWS via EKS
  • Collaborate with product and engineering teams for robust software development
  • Promote automation and self-service practices for operational efficiency
  • Lead initiatives aimed at achieving high availability and cost-effectiveness
  • Engage with customers and incorporate feedback into operations

Benefits

  • Meaningful stock options in a fast-growing company
  • 401(k) plan with 4% match
  • New HQ in Flatiron, NYC
  • Unlimited PTO
  • Health, dental, and vision insurance with 100% employee coverage
  • 12 weeks of paid parental leave
  • Commuter and fitness benefits
Full Job Description
About the role

We're looking for an Engineering Manager to lead a distributed SRE team to work on our production infrastructure. You will be responsible for supporting a team and driving a roadmap that makes Radar's services scalable and highly available worldwide. Radar is a high-throughput, data intensive application handling 1 billion+ API calls per day. Over the past year, Radar has been used on over 300MM devices worldwide. We run a multi-availability zone architecture and a major initiative is to enhance our deployment to be multi-region.

The stack:

We use Terraform to manage our infrastructure.

We deploy everything to AWS via EKS.

We use MongoDB deployed to Atlas.

We do CI and deployments via CircleCI.

We monitor production with CloudWatch, Grafana, Pingdom and PagerDuty.

DNS is managed by CloudFlare.

Most engineers are in the on-call rotation.

Our main server languages are TypeScript and Rust.

Our data pipelines are written in Airflow and Scala Spark.

We sponsor OpenStreetMaps, MapLibre and OpenAddresses.

How we use AI:
  • Engineers choose what AI tools they use, Claude and Codex being the most popular.
  • We're actively building Claude skills - for example we've taught it how to debug HorizonDB, our geospatial database.
  • All code changes are reviewed by an Engineer knowledgeable in that area. Claude and Codex also review all PRs.
  • There is a range of how much engineers use AI. Most use it daily if not weekly.
  • We are excited about what AI can do, but we also recognize the risks and don't compromise our coding standards.

How we work:

Most of our engineering team are former technical co-founders or former Radar interns from schools like Waterloo and CMU. Most engineers at Radar fit one of two molds, technically: either Staff level expertise in one stack, or Multi-Stack at any level. We say Multi-Stack because "Full-Stack" has the connotation of "Frontend and Backend", but Radar Engineers might also work on Mobile or Data engineering. Not that you need to be an expert in all of those, but a desire to learn, jump around to different stacks and get things done is the important part.

We care a lot about shipping fast and talking to customers. We're committed to our product vision of full-stack location infrastructure, but we also know that customer feedback is a treasure map to gold. Even though Slack is the brain of our company, working together in-person in our NYC HQ is the fastest way for us to get things done. We meet on Mondays to plan out work for the week in small groups and use Linear for planning. All projects are run by an Engineering lead, an executive and a Go-to-Market lead. Engineers figure out what to build, talk to customers, talk to prospects, help close them, get them live and make them successful.

One of the hardest and most valuable practices we have is Walk A Mile - which is shorthand for putting yourself in the user's shoes - but also for literally walking a mile and dogfooding the Radar SDK, because you can't create location infrastructure behind a desk - you have see how the device behaves in the real world. To us, a week is a long time, and we expect to ship big things every week.

The hiring process:

After a call with our Technical Recruiter, you'll do several technical Zoom calls with members of our engineering team: code screen, coding round, and system design round. If those go well we'll invite you to our NYC HQ for a final round interview. You'll meet one of our co-founders, someone from outside engineering, and meet more people from Radar. We'll go into more depth about how we work to see if there is a match.

What you'll do:
  • Manage, grow, and mentor a globally distributed SRE team
  • Be responsible for our observability, monitoring, and incident management practices
  • Work on core Radar infrastructure using Terraform all deployed to AWS via EKS
  • Work with our product, data, platform, and security engineers to ensure the development of highly-available, scalable, and reliable software
  • Champion automation and self-service capabilities over manual processes, and have grounded perspectives on AI-assised tooling
  • Drive critical company-level initiatives, for example, 99.99+% availability, multi-region deployment, and cloud cost-saving activities
  • Ensure compliance and security by default in our processes and infrastructure by partnering with our GRC and security teams
  • Have your work be used by 100's of millions of devices
  • Be part of the on-call rotation
  • Talk to Radar customers and prospects, hear their feedback, incorporate it into your work and make them successful

You should:
  • Be able to get your hands dirty and dive deep into production issues
  • Have experience managing a production AWS environment via Terraform
  • Have experience with high availability multi-region production infrastructure running on Kubernetes
  • Have experience managing large sharded Mongo clusters with heavy read-write workloads
  • Have experience at a high growth startup
  • Be interested in talking to customers or prospects and making them successful

Bonus points if you:
  • Are a former technical co-founder
  • Have experience with high throughput data intensive applications

You'll work with
  • Tim Julien, CTO
  • Leo Kim, VP Platform Engineering
  • Our Engineering, Customer Success, Sales Engineering and Sales teams
  • Our customers and prospects

What we offer:
  • Competitive salary
  • Meaningful stock options in a fast-growing company
  • 401(k) plan with 4% match
  • New HQ in Flatiron, NYC
  • Top-notch equipment
  • Catered lunches
  • Unlimited PTO
  • Health, dental, and vision insurance with 100% coverage for employees
  • 12 weeks of paid parental leave
  • Commuter and fitness benefits

We'll share full details of our benefits package at the offer stage. Benefits may vary by location.

About Radar.io

Radar.io is a location-based services company that provides a platform for businesses to build location-aware applications. The company's platform uses geofencing and other location-based technologies to provide businesses with real-time location data. Radar.io's platform can be used in a variety of industries, including retail, transportation, and hospitality.
Learn more about Radar.io
Size
50 employees
Industry
Founded
2014

Similar Jobs

More Jobs at Radar.io

More Information Technology Jobs

Find similar Engineering Manager, Site Reliability jobs: