Core FocusTo design, build, and maintain highly reliable data pipelines, storage systems, and transformational schemas; ensuring that secure, clean, and optimized data flows continuously across all core SaaS products and client reporting systems.
Roles & Responsibilities- ETL & Data Pipeline Engineering: Designing, constructing, and maintaining automated data pipelines (ETL/ELT) to ingest, clean, and transform disparate data sources into organized, high-performance repositories.
- Database & Schema Modeling: Designing logical and physical database schemas across relational (e.g., SQL Server, Azure SQL) and non-relational (e.g., Cosmos DB) platforms to support application scale and fast query execution.
- Data Warehousing & Architecture: Building and optimizing centralized data warehouses or staging environments that aggregate complex transactional insurance data for downstream reporting systems.
- Query Performance Tuning & Optimization: Monitoring, profiling, and tuning database performance - including query execution plan analysis, index optimization, and storage strategy adjustments.
- Data Security & Compliance Blueprinting: Implementing rigorous access controls, data masking, encryption standards, and retention policies to ensure absolute compliance with SOC 2 and insurance data-privacy rules.
Skills & Experience- Professional Core Experience: 3-5 years of dedicated experience operating as a Data Engineer or Database Developer managing complex, multi-source data environments (experience in high-compliance SaaS or financial services is a major plus).
- Advanced SQL & Procedural Scripting: Expert-level mastery of advanced SQL (including writing highly performant queries, complex joins, subqueries, and database optimization techniques).
- Data Pipe Programming: Solid experience writing scripts in Python or similar development languages to run automated API data extractions, cleansing routines, and custom integrations.
- Modern Cloud Infrastructure: Strong practical experience with cloud-native data services (e.g., Azure SQL, Cosmos DB, AWS Athena, or Snowflake) and cloud data pipeline engines (e.g., Azure Data Factory, dbt, or Airflow).
- Relational and Dimensional Modeling: Deep conceptual understanding of star schemas, snowflake schemas, and relational database normalization vs. denormalization strategies.
Success Metrics- Data Pipeline Uptime & Delivery: Maintain a greater than or equal to 99% success rate on scheduled ETL pipeline executions and automated data transfers.
- Query Response Time Baseline: Ensure key production databases maintain a target average query latency under 200ms for standard transactional read operations.
- Data Delivery Integrity Rate: Zero critical production incidents caused by data corruption, schema mismatches, or missing automated load steps per quarter.
- Support & Reporting Team Unblocked SLA: Resolve internally flagged database or schema pipeline blockages in an average of less than 4 hours to keep downstream business intelligence teams moving.