PDT Partners

Cloud Computing Engineer

PDT Partners • $195K — $225K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Engineering or Applied Sciences or equivalent professional experience.
  • 5+ years of experience in building and shipping software systems.
  • 2+ years of experience in running production compute platforms and scheduling systems.
  • Expertise in software engineering, including system design and debugging, with proficiency in languages like Python, Go, or Rust.
  • Experience with infrastructure-as-code tools like Terraform.
  • Strong written and verbal communication skills.
  • Hands-on experience managing inferencing and training hardware at scale.

Responsibilities

  • Design and implement cloud-based HPC systems, focusing on both engineering and operations.
  • Manage day-to-day operations of the HPC platform to ensure 24/7 availability.
  • Implement automation solutions across CI/CD pipelines and production metrics for efficiency.
  • Maintain a user-centric approach by collaborating with researchers and engineers to enhance HPC systems.
  • Perform capacity management and optimization of benchmarks for performance-critical workloads.

Benefits

  • Hybrid work environment requiring office attendance in New York City at least 3 days a week.
  • Collaborative flat team structure with diverse responsibilities spanning research and operations.
  • Support for professional development and growth within the organization.
  • Opportunities to work on cutting-edge cloud HPC projects in a fast-paced environment.
Full Job Description
The Cloud Computing team is a group of experts solving computing problems in the critical path of Research. We work directly with Research and Model Implementation teams and provide them with platform tools and computing resources to take their ideas from inception to real tradable products. We are looking for an ambitious and operationally minded software engineer to join our team as we mature and scale our cloud HPC platform to the next iteration of our firm-wide Research platform.

This is a hybrid position and will require the person to work from our New York City office at minimum 3 days a week.

Responsibilities:

We are a small flat team sitting at the cross-section of research, implementation, and platform infrastructure. Our team responsibilities span many areas. Including:
  • Design and implementation of cloud-based HPC systems. Our projects involve equal parts engineering and operations for success in our fast-moving environment. You will be expected to conceive and implement projects small and large in the intersection of HPC scheduling, metrics, containerization, software distribution, accelerated compute performance/efficiency, cloud architecture.
  • Running our HPC plant day-to-day. Our research environment is up 24/7, and we want to keep it that way. Everybody on the team contributes to the support of our platform, which thankfully is light because of our automation and quality work.
  • Implementing automation. We will always choose to work smart over working hard. You will be responsible for conception and implementation of automation from CI/CD pipelines to production metrics and monitoring of our cloud HPC platform.
  • Obsessive User Focus. All members of the team are expected to partner with researchers and engineers to deliver high-quality cloud HPC systems that are efficient and reliable. This includes leading projects to evolve it as our needs change.
  • Capacity management and benchmark optimization. Our demand for scale and performance is constant and involves challenging optimization problems for workloads critical to research and trading

Below is a list of skills and experiences we think are relevant. Even if you don't think you're a perfect match, we still encourage you to apply because we are committed to developing our people.
  • Bachelors or Masters degree in an Engineering or Applied Sciences field from a rigorous academic program or equivalent professional experience.
  • 5+ years of experience building and shipping software systems
  • 2+ years running production compute platforms and scheduling systems.
  • Mastery of core software engineering concepts such as system design, testing, debugging, building reliable production systems with fluency in a language such as Python, Go, Rust etc.
  • Experience with a cloud-based infrastructure-as-code tool such as Terraform
  • Excellent written and verbal communication skills
  • Experience working with or supporting researchers and/or other developers is a plus
  • Hands-on experience managing inferencing and training hardware at scale
  • Knowledge of NVIDIA GPU management, Kubernetes, Slurm, and the wider large-scale compute ecosystem

The salary range for this role is between $195,000 and $225,000. This range is not inclusive of any potential bonus amounts. Factors that may impact the agreed upon salary within the range for a particular candidate include years of experience, level of education obtained, skill set, and other external factors.

About PDT Partners

PDT Partners is a quantitative investment firm that uses advanced mathematical models and algorithms to identify and exploit market inefficiencies. The company was founded in 1993 by Peter Muller, a former head of Morgan Stanley's quantitative trading group. PDT Partners has a strong track record of generating alpha for its investors, and has consistently outperformed its benchmarks over the long term. The company's investment strategies are based on a deep understanding of market dynamics and a rigorous approach to risk management.
Learn more about PDT Partners
Size
200 employees
Industry

Similar Jobs

More Jobs at PDT Partners

More Information Technology Jobs

Find similar Cloud Computing Engineer jobs: