We are seeking a highly technical
Principal Data Engineer to design, scale, and evolve our enterprise data ecosystem. In this role, you will serve as the technical authority for our data infrastructure, driving the strategy for data warehouse design, real-time streaming, and robust governance.
As a Principal Engineer, you will balance hands-on technical leadership with strategic cross-functional collaboration, mentoring engineers, and aligning our data capabilities with long-term business goals. You are the strategic anchor and technical authority for our enterprise data ecosystem. In this role, you will design, govern, and optimize our Microsoft Fabric and Power BI design to transform complex, multi-modal data streams-including near-real-time telematics, operational routing, and financial ERP systems-into secure, scalable, and actionable intelligence.
You will bridge the gap between technical execution and business strategy, ensuring our data platform is robust, compliant with federal regulations (FERPA), and built for high-performance self-service analytics.
ResponsibilitiesEnterprise Data Design- Define Medallion Standards: Author, document, and defend the design boundaries and ingestion logic across the Bronze, Silver, and Gold layers within Fabric.
- Tool & Pattern Selection: Establish clear organizational frameworks for when to deploy specific Fabric components (e.g., Mirroring vs. Dataflow Gen2 vs. Data Factory Pipelines).
- Power BI Storage Strategy: Standardize dataset configurations, defining precise criteria for utilizing DirectLake, Import, or DirectQuery modes to balance cost, freshness, and performance.
Integration & Advanced Pipeline Engineering- Multi-Speed Ingestion Patterns: Design repeatable, hardened ingestion patterns for diverse data velocities:
- Near Real-Time: GPS and IoT telematics streams.
- Operational Batch: Dispatch, scheduling, and routing systems.
- Transactional/Financial: Enterprise ERP data.
- Resiliency & Frameworks: Own the end-to-end framework for data ingestion into OneLake, embedding enterprise-grade error handling, automated retry logic, and proactive monitoring/alerting systems.
Data Modeling & Master Data Management (MDM)- Gold Layer Definition: Design enterprise-wide dimensional models and star schemas within the Gold layer to serve as the single source of truth.
- Semantic Layer Standardization: Establish strict governance for Power BI semantic models, including DAX measure naming conventions, hierarchy design, and Row-Level Security (RLS) structures.
- Master Data Management: Spearhead the definition and enforcement of canonical master data entities across disparate systems, specifically governing unified definitions for Vehicle IDs, Driver IDs, and Student Records.
Data Governance, Security, and Compliance- Access Control & Lifecycle: Enforce Fabric workspace security utilizing Microsoft Entra ID / Security Groups for identity management, eliminating ad-hoc individual user provisioning.
- Data Protection & Auditability: Implement Microsoft Purview sensitivity labels, robust Row-Level Security (RLS), and comprehensive audit trails, serving as the primary technical point of contact for data security audits.
Business Enablement & Analytics Delivery- Requirements Translation: Act as the primary conduit between operations, safety, finance, and district stakeholders, translating complex operational requests (e.g., "On-time arrival rate by district") into technical data products.
- Productized Delivery: Lead the deployment of certified, user-friendly Power BI apps and dashboards, reducing friction for business users and eliminating the need for stakeholders to write direct SQL/KQL queries against the Gold layer.
Technical Profile & Qualifications- Platform Expertise: Deep, hands-on experience with Microsoft Fabric, OneLake, Power BI, and the Azure Data stack.
- Data Modeling Mastery: Proven track record building enterprise Star Schemas, handling Slowly Changing Dimensions (SCDs), and establishing Master Data Management (MDM) protocols.
- Security & Compliance: Direct experience securing highly sensitive data environments; familiarity with FERPA, HIPAA, or similar strict privacy frameworks is highly preferred.
- Communication & Documentation: Exceptional ability to articulate complex technical decisions to executive leadership, document blueprints, and mentor engineering teams.
The salary is $140k - $160k annually.
QualificationsQualifications:- Experience: 10+ years of experience in data engineering, with at least 2+ years of hands-on experience designing enterprise solutions within Microsoft Fabric or advanced Azure Synapse/Databricks environments.
- Fabric Ecosystem Mastery: Deep technical understanding of Fabric capacities, OneLake shortcuts, Managed Private Endpoints, and the interplay between Lakehouses and Warehouses.
- Advanced Spark & Delta Lake: Expert-level knowledge of Apache Spark (PySpark) and the Delta Lake storage format, including optimization techniques like V-Order, Z-Order indexing, and liquid clustering.
- Languages: Expert in SQL and Python (including libraries like Pandas, Polars, or PySpark).
- Orchestration & DevOps: Experience with Fabric Git integration, Azure DevOps/GitHub Actions, and Infrastructure as Code (Terraform) for Fabric resource provisioning.
We offer medical, dental, vision, basic life insurance coverage, holiday pay, and PTO accrual. Additionally, employees are able to enroll in a retirement savings plan.