Infrastructure and Data Engineer
Meet the TeamTechnology is core to the health and growth of Millennium's business. The firm's active, multi-manager business model demands flexible, scalable technology and advanced proprietary systems, including the development of next-generation analytical and trading capabilities. The Data Science team builds and supports data infrastructure, applications, and AI and machine learning systems used by people and platforms across the business.
What You'll Do- Build and maintain Python and SQL data pipelines that support feature stores, embeddings, and model training and inference workflows.
- Develop infrastructure for extracting, transforming, and loading data from Snowflake, SQL Server, streaming sources, and other systems across on-premises and cloud environments.
- Design, deploy, and maintain reproducible, scalable environments using infrastructure-as-code tools such as Terraform and CloudFormation.
- Manage region-redundant Airflow orchestration, Docker and Kubernetes services, and the NGINX, Gunicorn, and Django web-serving stack.
- Strengthen user-facing applications and APIs by improving authorization, load balancing, containerization, and CI/CD deployment pipelines.
- Build production infrastructure for LLM-based applications, including retrieval-augmented generation pipelines, vector databases, embeddings, and third-party or self-hosted model APIs.
- Implement monitoring, logging, observability, and analytics across data pipelines, application infrastructure, and model-serving endpoints to provide insight into system health, usage, and performance.
- Drive improvements across the technology stack by automating manual processes, strengthening data governance and access controls, improving scalability, and reducing infrastructure and inference costs.
What You Bring- Two or more years of professional experience with a master's degree, or three or more years with a bachelor's degree, in Computer Science, Statistics, Informatics, Information Systems, or another quantitative field.
- Advanced Python skills, including experience with Pandas, NumPy, and SciPy, as well as familiarity with machine learning and AI libraries such as PyTorch, scikit-learn, or orchestration frameworks such as LangChain.
- Strong SQL and data engineering experience, including relational databases such as Microsoft SQL Server or PostgreSQL, modern cloud data warehouses such as Snowflake, and pipeline orchestration with Airflow.
- Hands-on experience with AWS services, Docker, infrastructure-as-code tools such as Terraform or CloudFormation, and preferably Kubernetes.
- Practical knowledge of production AI and MLOps systems, including vector databases, embedding-based retrieval, LLM APIs, prompt engineering, context management, RAG architecture, model versioning, feature stores, and model monitoring.
- Strong computer science and infrastructure fundamentals, including distributed systems, data structures, threading, memory management, Unix/Linux environments, CI/CD, security, and observability.
- Excellent written and verbal communication skills, with the ability to work independently and collaboratively while managing multiple priorities in a fast-paced environment.
- A proactive, detail-oriented approach to problem-solving, with ownership of outcomes and the ability to deliver quality code as technologies and priorities evolve; knowledge of financial instruments is highly valued.
The estimated base salary range for this position is $175,000 to $250,000, which is specific to New York and may change in the future. Millennium pays a total compensation package which includes a base salary, discretionary performance bonus, and a comprehensive benefits package. When finalizing an offer, we take into consideration an individual's experience level and the qualifications they bring to the role to formulate a competitive total compensation package.