MongoDB's Replication Team builds the infrastructure that enables high availability, fault tolerance, automatic failover, and tunable consistency. As an engineer on the team, you will design and implement distributed-systems features that protect data and keep applications available under demanding operating conditions.
You will work primarily in C++ on core database code, partner with engineers across MongoDB, and help shape features that are central to major MongoDB releases. This is an opportunity to apply distributed-systems fundamentals to a widely used database while solving challenging problems in correctness, performance, and operability.
We're looking to speak with candidates based in New York City for our in-office working model.
What you will do- Design and implement replication features based on the Raft consensus protocol
- Improve failover behavior, availability, correctness, and performance across the replication system
- Write production-quality C++ and the unit, integration, and system tests needed to demonstrate correctness
- Use JavaScript and Python where appropriate to extend test coverage and validate end-to-end behavior
- Diagnose test failures, investigate bugs, and drive issues through root-cause analysis and resolution
- Measure the performance impact of code changes and prevent or resolve regressions
- Collaborate with partner engineering teams and stakeholders on large, cross-functional initiatives
- Investigate distributed-systems issues raised by customers and Technical Support, communicate findings clearly, and help deliver durable fixes
- Participate in code reviews, design reviews, and technical discussions that improve the quality of the team's work
- Interview candidates and mentor junior engineers and interns
What you bringRequired
- At least five years of experience programming, debugging, and performance-tuning distributed or highly concurrent software systems
- Strong systems fundamentals, including multithreaded programming, concurrency, debugging, and performance profiling
- Practical experience with one or more distributed-systems concepts, such as consensus protocols, replication, distributed transactions, or fault tolerance
- Experience working with complex production codebases and diagnosing problems that cross subsystem boundaries
- Strong written and verbal technical communication skills
- A collaborative approach and the ability to make realistic assessments of project complexity and delivery risk
Preferred
- Professional experience developing production software in C++
- Familiarity with database internals or core components of data-processing systems
- Experience with Raft or another consensus protocol
- Experience writing unit, integration, and end-to-end tests in C++, JavaScript, or Python
- Experience responding to customer-impacting production issues in distributed systems
- Experience mentoring engineers or interns and contributing to technical hiring
Candidates with additional relevant experience may be considered for more senior roles.
What success looks likeIn your first month
- Build a working understanding of MongoDB's replication architecture, development workflow, and testing strategy
- Begin resolving scoped bugs and contributing effectively to code reviews
In your first three months
- Contribute production C++ code to a project targeted for an upcoming MongoDB release
- Diagnose and resolve issues identified by testing, customers, or partner teams
In your first six months
- Take regular ownership of code reviews and participate meaningfully in feature design reviews
- Independently drive scoped technical work from investigation through implementation and validation
In your first twelve months
- Lead the development of a new feature or substantial improvement from design through delivery
- Help mentor new engineers and contribute to the team's technical direction
MongoDB's base salary range for this role in the U.S. is:
$106,000-$209,000 USD