We're seeking a motivated Data Engineer to join our growing team. You'll be responsible for building and maintaining data pipelines and ensuring our solutions have access to clean, reliable, and well-structured data. This is an excellent opportunity to learn and grow in a fast-paced environment where you'll work alongside other data engineers (incl. software developers), and product teams. You'll gain hands-on experience with modern data technologies.
Key ResponsibilitiesData Pipeline Development- Build and maintain ETL/ELT pipelines to ingest data from various sources (databases, APIs, files, third-party systems)
- Transform raw data into structured formats suitable for application use
- Implement data validation and quality checks throughout pipelines
- Schedule and monitor automated data workflows
- Debug and fix pipeline failures promptly
Data Integration & API Support- Build integrations between different data systems and applications
- Create and maintain APIs for data access and retrieval
- Support agent tool integrations that require data access
- Work with external APIs to fetch and sync data
- Document data schemas, pipelines, and integration points
Data Quality & Monitoring- Implement data quality checks and validation rules
- Monitor data pipeline health and set up alerts for failures
- Investigate and resolve data inconsistencies or anomalies
- Create data quality reports and dashboards
- Maintain data lineage and documentation
Collaboration & Support- Work closely with other engineers to understand data requirements for agentic workflows
- Support product and engineering teams with data-related questions
- Participate in sprint planning and Agile ceremonies
- Contribute to technical discussions and code reviews
- Document processes, pipelines, and best practices
Required QualificationsEducation & Experience- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field
- Experience with SQL and relational databases
- Basic understanding of data pipeline concepts and ETL processes
- Exposure to cloud platforms (AWS, Azure, or GCP)
Technical Skills- Strong proficiency in Python for data processing and scripting
- Solid SQL skills - writing queries, joins, aggregations, and optimizations
- Experience with at least one relational database (PostgreSQL, MySQL, SQL Server)
- Understanding of data modeling concepts (normalization, star schema, etc.)
- Familiarity with version control using Git
- Basic understanding of Linux/Unix command line
- Knowledge of data formats (JSON, CSV, Parquet, etc.)
Preferred Skills- Experience with Python data libraries: polars, pandas, numpy
- Familiarity with ETL/orchestration tools: Airflow, Prefect, Dagster, or similar
- Basic understanding of APIs and REST principles
- Knowledge of containerization (Docker)
- Understanding of data warehousing concepts
Soft Skills- Eager learner - enthusiastic about learning new technologies and best practices
- Problem solver - logical approach to debugging and troubleshooting
- Detail-oriented - careful with data quality and accuracy
- Collaborative - works well in team environments and asks for help when needed
- Communicator - can explain technical concepts clearly
- Self-motivated - takes initiative and ownership of tasks
- Adaptable - comfortable with changing priorities in an agile environment
What You'll Work WithProgramming Languages- Python (primary) - polars, requests, SQLAlchemy (or other ORMs)
- SQL (extensive use across multiple databases)
- Understanding of Bash scripting for automation
- Understanding of containers (e.g. Docker)