Job Summary
We are seeking a Python Developer with strong data engineering expertise to design, develop, and maintain high-performance data processing pipelines using modern Python frameworks and tools. The role will involve working with large-scale datasets, containerized applications, distributed computing platforms, analytical databases, and event-driven architectures to deliver scalable and reliable data solutions. The position is based in Toronto, ON, with 4 days onsite.
Key Responsibilities
• Develop and optimize data manipulation workflows using pandas and polars for efficient processing of large datasets.
• Design and implement containerized applications using Docker and Kubernetes to support scalable and reliable deployments.
• Build and maintain data pipelines integrating with ClickHouse columnar databases for analytical workloads.
• Develop event-driven architectures using NATS messaging systems for asynchronous data processing.
• Implement distributed computing solutions using Dask for processing datasets beyond single-machine memory constraints.
• Write comprehensive unit and integration tests using pytest and maintain appropriate code coverage standards.
• Manage source code using Git and follow established branching and collaborative development practices.
• Optimize data processing code and workflows for performance and scalability.
Required Qualifications
• Advanced proficiency in Python, pandas, and polars for data manipulation, transformation, and analysis.
• Experience optimizing Python code and data processing workflows for large datasets.
• Hands-on experience with Docker, including building container images and composing multi-container applications.
• Knowledge of Kubernetes for container orchestration and deployment management.
• Working knowledge of ClickHouse or similar columnar databases for OLAP workloads and analytical queries.
• Familiarity with NATS.io for message-driven systems and asynchronous workflows.
• Proficiency with pytest for unit testing, integration testing, and maintaining code coverage.
• Experience with Dask for parallel processing and out-of-core computations.
• Strong command of Git workflows, branching strategies, and collaborative development practices.
Preferred Qualifications
• Experience with additional Python libraries for data science and machine learning.
• Familiarity with CI/CD pipelines and DevOps practices.
• Background in financial services or capital markets data systems.