Cars Commerce
• $118K — $148K *Qualifications
Responsibilities
Benefits
About the Role -
[Remote and must be located in the Chicago area. May require onsite interviews]
Join us in shaping a future of Automotive Commerce. Cars Commerce builds solutions for automotive dealerships, manufacturers, and consumers. Our driving force is to deliver a single platform that simplifies everything about buying and selling cars.
Our Data and AI Platform team is embarking on a transformational journey, building innovative and integrated services that redefine how we deliver value to our customers and consumers alike. This is a rare opportunity to join a data-first company at a genuine inflection point with strong leadership, a clear vision, and an enormous amount of meaningful work to be done. For a data professional who wants to build, create, and make a real impact this is the place and the moment.
About You -
As a Senior Data Engineer on our Data Platform Engineering Team, you will help revitalize our Data Services and tooling across our platform with a comprehensive technical strategy to simplify and build scalable systems. This is an exceptional opportunity to shape the future of our platforms and drive meaningful change at scale, making a lasting impact within the automotive industry.
As a Senior Data Engineer, you own a domain of our data platform end to end not just implementing pipelines, but designing them and making the call on ambiguous, harder problems with little hand-holding. You set the quality bar for your area: data contracts with upstream producers, automated testing and anomaly detection, and monitoring that catches issues before the business does.
You9ll build and tune production-grade pipelines using PySpark and advanced SQL, and design the dimensional models and medallion-style (bronze/silver/gold) data layers that Data Science, BI, and Analytics teams build on top of. You9re also accountable for the performance and reliability of the warehouse assets your domain produces, including query tuning and cost-aware design in Redshift.
This role is also where we9re pushing the platform forward: using Al-assisted development as a core part of how the team builds and ships, and helping lay the groundwork for a knowledge graph that connects our core entities and metrics across the business.
Why You9ll Be Excited About This Role
This opportunity enables you to...
Design, build, and maintain robust production pipelines and reusable data assets for analytics, Data Science, and BI, using PySpark and SQL-based transformation patterns.
Own data quality and reliability for your domain through data contracts with upstream producers, automated data quality checks, and anomaly detection.
Tune PySpark jobs and pipeline performance for runtime, memory, partitioning, and compute cost at production scale.
Design and maintain dimensional models and medallion-style (bronze/silver/gold) data layering, producing reusable, well-documented assets for Data Science and BI teams.
Optimize Redshift performance; query tuning, Spectrum external tables, Glue Data Catalog, and materialized views for the warehouse assets your domain owns.
Own warehouse pipeline SLAs, freshness monitoring, and incident response for your domain; participate in on-call.
Break down ambiguous business problems and resolve them, mapping day-to-day work to broader increment goals.
Drive adoption of Al-assisted development practices across the team - using tools like Claude Code, Cursor, or Amazon Q to accelerate pipeline development, testing, and documentation, not just code completion.
Partner directly with stakeholders on requirements and prioritization; establish and document engineering standards for your area.
Review code and ensure the quality of what the team ships
Mentor and coach more junior engineers, leveling up their technical judgment and ownership over time.
We9re Excited About You Because...
5+ years building and maintaining production data pipelines, with demonstrated ownership of a domain, a significant migration, or an architecture decision.
Proficient in PySpark and Python, including performance tuning for runtime, memory, and cost.
Advanced SQL for complex data manipulation, joins, aggregation, and window functions.
Solid cloud experience; AWS strongly preferred.
Strong data modeling skills; dimensional design and grain, and medallion-style layering
Delta Lake/Iceberg fluency table format internals, compaction, vacuuming, and schema evolution.
Experience with Redshift query optimization, Spectrum external tables, Glue Data Catalog, and materialized views.
Experience establishing data contracts and schema validation with upstream producers, and working with observability tooling (e.g., Great Expectations, Metaplane, Monte Carlo, or similar).
Proficient with Airflow (or similar orchestration)
API development for consuming and serving data or model outputs.
Comfort with infrastructure-as-code and mature CI/CD practices (e.g., GitHub Actions, Jenkins).
Hands-on experience with AI-assisted development tools (e.g., Claude Code, Cursor, or Amazon Q).
Interest in or experience with knowledge graphs and graph-based data modeling you9ll help design and build out this capability as a core part of our semantic layer and catalog work.
Demonstrates ability to mentor without being the go-to answer key you teach people how to reason through a problem rather than just handing them the solution.
Bachelor9s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Bonus:
Experience with streaming/CDC patterns (Kafka, MSK, or Flink).
Exposure to data cataloging and lineage tools (Atlan or similar).
Experience in a high-availability, distributed environment.
AWS certifications.
Our Comprehensive Benefits Package includes:
Medical, Dental & Vision Healthcare Plans
New Hire Stipend for Home Office Set-Up
Generous PTO
Paid Holidays, Floating Holiday, Volunteer Day, Recharge Day
Learn more about our Benefits, Perks, & Culture on our LinkedIn Life Pages!
Similar Jobs





More Jobs at Cars Commerce
More Information Technology Jobs