The Position:
We're seeking a Data Engineer to join a new Data team building the pipelines and datasets that turn telematics and operational data from MacLean's vehicles into trusted analytics and customer-facing products. Reporting to the Team Lead - Data, you'll be a hands-on contributor on the migration from our legacy telemetry platform to a modern cloud data warehouse, and on the streaming pipeline that lands next-generation field data from connected vehicles.
You'll work across the full ingestion-to-consumption path: change-data-capture from legacy relational systems into the cloud warehouse, streaming consumers landing high-frequency vehicle data, and curated transforms that unify historical, batch, and streaming sources into a single schema defined by our vehicle data semantics contract. Downstream, your datasets power MacLean's commercial telemetry products: fleet utilization reporting, fault and signal-health alerting, OEM KPI dashboards, and a roadmap toward predictive maintenance and anomaly detection.
This role sits within the Telematics, Automation, and Controls Division and partners closely with the Platform team, the Vehicle Interface team, and product stakeholders. You'll write production Go and SQL day to day, operate at the intersection of industrial IoT and modern data platforms, and ship across MacLean's 32+ vehicle types and global footprint.
Responsibilities and Duties:
- Build and operate ingestion pipelines from vehicles and legacy systems into the cloud warehouse - including change-data-capture from relational sources, streaming consumers writing high-frequency telemetry, and refactoring existing batch ingestion onto the new platform.
- Develop and maintain curated transforms that unify historical, batch, and streaming telemetry sources into a single schema aligned to our vehicle data semantics contract.
- Model data in the warehouse across raw, curated, and semantic layers, and build the certified datasets that downstream dashboards, alerts, and analytics products consume.
- Build the data flows behind MacLean's commercial telemetry products - utilization reporting, hour meter-based dashboards, fault and signal-health alerting, OEM KPI dashboards, operator and shift analysis, and near-real-time data delivery for streaming use cases.
- Implement schema management, data contracts, data quality checks, and SLA monitoring across pipelines and datasets, including dual-read and dual-write validation during migration cutovers.
- Write production Go services and SQL transforms; contribute Python where appropriate for analytics tooling and ML feature pipelines.
- Partner with the Platform team on the messaging subject namespace, message envelope standards, and stream provisioning - consuming the platform's governance model rather than defining it, and surfacing data-side requirements back to it.
- Partner with the Vehicle Interface team on signal definitions, since the vehicle data semantics contract is the structural source of truth for curated datasets.
- Support the data science workstream by building reliable feature datasets for predictive maintenance, root cause diagnostics, and anomaly detection, and by collaborating on production model pipelines.
- Contribute to the evaluation and integration of a modern observability and analytics platform against curated datasets, and help migrate existing portal use cases onto it.
- Follow team engineering standards for code quality, testing, CI/CD, containerized deployment (Docker/Kubernetes), and release governance for pipelines, datasets, and models.
Qualifications: The ideal candidate will possess strong communication skills, leadership abilities, and problem-solving capabilities. They will demonstrate organizational skills, discipline, an aptitude for proactive learning, and a positive attitude. Additionally, they should meet the following qualifications:
- Minimum 3 years of experience in data engineering or a closely related discipline, building and operating production data pipelines.
- Bachelor's degree or higher in computer science, software engineering, data engineering, or a related field.
- Strong programming skills in Go and SQL; proficiency in Python is an asset.
- Hands-on experience with at least one modern cloud data warehouse - Snowflake preferred - and with multi-layer data modeling (raw → curated → semantic, or Bronze/Silver/Gold)..
- Experience building ingestion or streaming pipelines using technologies such as Kafka, RabbitMQ, 0MQ, or NATS. Familiarity with NATS JetStream and the Snowflake Streaming SDK is an asset
- Working knowledge of industrial or IoT protocols such as OPC UA and MQTT (including Sparkplug B), and a willingness to ramp on the OPC UA information model.
First Year Goals:
- Develop a strong understanding of the current and future state of MacLean Engineering's cloud data ecosystem, including data sources from connected underground mining equipment, backend systems, and existing analytics workflows.
- Navigate the learning curve to align with the team's development pace, internal tools, and operational processes for data management, validation, and deployment.
- Gain deep familiarity with existing data pipeline, from on-machine telemetry to cloud ingestion, storage, and reporting layers, while identifying opportunities to improve performance and consistency.
- Collaborate with software engineers to redesign and optimize backend structures and APIs to enable more efficient, accurate, and scalable data querying and analysis.
- Contribute to grouping raw data into logical "data bricks" that support robust data modeling, contextualization, and canonical frameworks, improving downstream analytics and dashboard efficiency.
- Support the transition of legacy scripts and code into more scalable, maintainable, and high-performing solutions aligned with data best practices.
- Begin developing reusable datasets and validated pipelines that serve as reliable foundations for dashboards, performance metrics, and advanced analytics.
- Participate in query and code reviews, documentation efforts, and validation processes to strengthen data reliability and team-wide analytical standards.
- Practical experience with CDC, schema management, data contracts, and data quality frameworks.
- Experience with containerized deployment (Docker, Kubernetes) and CI/CD for data pipelines.
- Working knowledge of one or more BI or observability tools (e.g., Power BI, Grafana, Observe) sufficient to build and validate dashboards on top of curated datasets.
- Strong verbal and written communication skills in English.
- Proficiency in Spanish is an asset given cross-site collaboration with our Queretaro, Mexico office.
- Experience in mining, industrial, or heavy-equipment telematics is an asset.
- Exposure to MLOps practices or production ML feature pipelines is an asset.
This position will be located in any of our Sudbury, Collingwood, Barrie or Waterloo Ontario; Querétaro, Mexico