Role DescriptionSummaryWe are seeking an Associate Data Engineer to join the Trade Data Repository (TDR) team within Market Surveillance Technology. The role will support the development, enhancement, monitoring, and modernization of a critical surveillance data platform responsible for ingesting, validating, transforming, and distributing trade and order data to surveillance systems and analytical applications. TDR currently operates on Azure SQL Server and Azure Blob Storage and is actively migrating to an Azure Databricks Lakehouse architecture utilizing Bronze, Silver, and Gold data layers.
The successful candidate will work closely with senior engineers, business analysts, surveillance stakeholders, Data Governance teams, and upstream application teams to build reliable and scalable data pipelines supporting regulatory and compliance objectives.
Role Objectives: DeliveryKey ResponsibilitiesData Engineering & Development- Develop and maintain SQL-based ingestion, transformation, and reconciliation processes within the TDR platform.
- Design and optimize complex SQL queries, stored procedures, views, functions, and ETL processes.
- Build and support Azure Data Factory (ADF) data pipelines for ingestion of structured and semi-structured data.
- Develop Databricks notebooks using PySpark, Spark SQL, and Python.
- Participate in migration of TDR workloads from Azure SQL Server to Azure Databricks Lakehouse.
- Support implementation of Bronze (Pre-Stage), Silver (Stage), and Gold (Outbound) data layers.
Azure Platform Development- Develop integrations involving:
- Azure SQL Database
- Azure Blob Storage / ADLS Gen2
- Azure Data Factory
- Azure Databricks
- Delta Lake
- Build reusable ingestion and transformation frameworks.
- Support batch processing, incremental loads, and historical data processing.
- Implement metadata-driven pipeline designs and parameterized frameworks.
Data Quality & Governance- Implement data quality validations and reconciliation controls.
- Develop source-to-target validation checks and record-count reconciliations.
- Support KDE (Key Data Element) controls and exception management processes.
- Participate in root-cause analysis of data quality issues.
- Assist in maintaining auditability, lineage, and data governance controls required for regulatory compliance.
Production Support & Monitoring- Monitor daily processing and data loads.
- Diagnose and resolve data pipeline failures.
- Support incident management, defect remediation, and change implementation.
- Perform performance tuning for SQL and Databricks workloads.
- Partner with business and support teams during production incidents.
Regulatory & Surveillance Data Processing- Support ingestion and processing of:
- Equities trades and orders
- Futures trades and orders
- Fixed Income trades
- FX trades
- Interest Rate Swaps
- Employee trades
- Surveillance alerts and analytics outputs
- Ensure data delivered to surveillance platforms is accurate, timely, and complete.
SDLC, Agile & DevOps- Participate in Agile delivery processes including sprint planning, daily stand-ups, backlog refinement, retrospectives, and release planning.
- Utilize Jira for user story management, defect tracking, sprint execution, and workflow management.
- Utilize Confluence for technical documentation, architecture diagrams, knowledge sharing, operational procedures, and project collaboration.
- Perform unit testing and support SIT, UAT, regression testing, and production validation activities.
- Support CI/CD deployment processes across Development, QA, UAT, and Production environments.
- Utilize Git and Azure DevOps repositories for source control, branch management, code reviews, and release management.
- Collaborate with developers, business analysts, QA teams, and stakeholders throughout the SDLC.
- Prepare and maintain technical documentation, source-to-target mappings, support procedures, and operational run-books.
- Participate in production support, incident management, root cause analysis, and continuous process improvement initiatives.
Qualifications and SkillsRequired QualificationsEducation- Bachelor's degree in Computer Science, Information Systems, Engineering, Data Analytics, or related discipline.
Experience- 2-5 years of experience in Data Engineering, ETL Development, or Data Warehousing.
- Experience developing solutions in Microsoft Azure environments.
- Experience working with financial services, capital markets, risk, compliance, or surveillance data preferred.
Required Technical SkillsSQL & Database Development- Advanced SQL development
- Stored Procedures
- Views
- Performance tuning
- Query optimization
- Data modeling
- Data reconciliation
Azure Data Engineering- Azure SQL Database
- Azure Data Factory (ADF)
- Azure Blob Storage / ADLS Gen2
- Azure Databricks
- Delta Lake
- PySpark
- Spark SQL
- Python
Data Integration- ETL/ELT Design
- Batch Processing
- Incremental Processing
- Source-to-Target Mapping
- Data Validation
- Data Reconciliation
DevOps- Git
- Azure DevOps
- CI/CD concepts
- Release Management
Preferred Qualifications- Experience with Market Surveillance, Trade Surveillance, or Regulatory Reporting.
- Familiarity with Actimize, Quantexa, Compliance Analytics, or surveillance ecosystems.
- Experience implementing Delta Lake architectures.
- Exposure to Unity Catalog and data governance frameworks.
- Understanding of financial products including Equities, Fixed Income, Futures, FX, and Derivatives.
- Experience working within regulated financial institutions.
SMBC's employees participate in a Hybrid workforce model that provides employees with an opportunity to work from home, as well as, from an SMBC office. SMBC requires that employees live within a reasonable commuting distance of their office location. Prospective candidates will learn more about their specific hybrid work schedule during their interview process. Hybrid work may not be permitted for certain roles, including, for example, certain FINRA-registered roles for which in-office attendance for the entire workweek is required.