Full Job Description
The Senior Data Engineer designs, develops, and maintains secure, scalable cloud-based data architectures and data pipelines. The position supports advanced analytics, fraud detection, machine learning, and investigative operations by delivering reliable, high-quality data solutions within the Microsoft Azure environment.
Key Responsibilities:
• Develop new datasets to support loan fraud identification, audits, and investigations.
• Optimizes the performance of common queries in the database.
• Creates and implements data models to improve the efficiency and utility of the database.
• Proposes and implements architectural improvements to reduce cloud costs.
• Develops, optimizes, and maintains ELT/ETL pipelines
• Monitors pipelines to ensure performance and regular updating of datasets.
• Works with a Cloud Architect to develop new ways of accessing and transporting data.
• Aids in the development of dashboards or reports to visualize data and storytelling.
• Assists in ingestion, transformation, and configuration of data in cloud-based data lake or data warehouse.
• Assists with exploration, structuring, and analysis of unstructured and semi-structured data from SQL and data lake environments.
• In collaboration with data architecture staff, authors SOPs governing creation, maintenance, and monitoring of all assets in Azure environment.
• Design, implement, and maintain ELT/ETL pipelines for efficient processing of source data in Azure Synapse and Azure Machine Learning (using SDK V1 and SDK V2).
Required:
• Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field; OR 5 years of applied work experience in any of the same fields.
• Minimum 5 years of experience in Azure DevOps, Azure Synapse and Azure Machine Learning.
• Five (5) years of hands-on experience in each of the following:
a. Maintaining SQL databases and conducting advanced operations in SQL and T-SQL.
b. Designing, implementing, and maintaining ELT/ETL processes in cloud-based data analytics environments.
• Three (3) years of hands-on experience in each of the following:
A. Working in Azure Synapse and Azure Machine Learning, with the modern data stack. Certifications equivalent to DP-203 or DP-900.
B. Manipulating data in Python. Pandas required. PySpark/Polars preferred. Experience developing reusable, modular code preferred.
Preferred:
• Implementing pipelines and infrastructure using code-first approaches (Python SDK, CLI, REST APIs, or IaC tooling)
• Additional certifications preferred: AZ-305
• Implementing source control and CI/CD workflows
• Demonstrated familiarity with AI coding assistants and LLM integration patterns
• Experience with REST APIs, Infrastructure as Code (IaC), and Azure DevOps.
• Experience supporting AI, Large Language Models (LLMs), and Natural Language Processing (NLP).
• Experience supporting Federal Government or Inspector General organizations.