Job Overview
Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive, and iterative delivery environment? At Mercedes-Benz USA, you'll be part of a group who love to solve real problems and meet real customer needs
We are seeking an experienced Senior Data Engineer who will be part of a team building a robust data platform enabling end-to-end capabilities for Reporting, Machine Learning, Generative AI, and Agent-based AI products supporting a variety of users including data engineers, analysts, scientists, AI engineers, and internal/external business partners
The Principal Data Engineer is a senior technical leader responsible for defining complex problem spaces, setting architectural direction, and delivering scalable, enterprise-grade data platforms and products. This role operates effectively in high-ambiguity environments, owns outcomes and business impact, and establishes standards and frameworks adopted across multiple teams and domains. The Principal DE plays a critical role in bridging traditional data engineering with modern AI/GenAI data infrastructure, ensuring the data platform evolves to serve both analytics and AI workloads at enterprise scale
Responsibilities
Data Platform Architecture & Engineering (40%)
- Define and drive enterprise data engineering architecture, standards, and best practices across the organization.
- Design and deliver scalable, high-performance data platforms supporting analytics, machine learning, and generative AI use cases.
- Architect data lake, warehouse, and lakehouse solutions using Databricks, Delta Lake, and Unity Catalog.
- Design and deliver data platform capabilities that support generative AI and agent-based workloads, including embedding pipelines, vector store integration, and knowledge base management.
- Identify systemic gaps in data quality, reliability, performance, and cost, and drive solutions end-to-end.
- Establish reusable frameworks, patterns, and playbooks for data pipeline development adopted across teams.
- Define data partitioning, optimization, and caching strategies for high-volume and low-latency workloads.
- Lead technical design reviews and architecture decision records (ADRs) for major data platform initiatives
AI/GenAI Data Infrastructure (20%)
- Design and implement data architectures that serve AI/ML model training, RAG pipelines, and agent-based systems.
- Build and govern embedding pipelines, vector database ingestion workflows, and knowledge base refresh processes.
- Define data quality, freshness, and governance standards for AI-consumed datasets.
- Architect unstructured data processing pipelines (document parsing, chunking strategies, metadata enrichment) for GenAI consumption.
- Collaborate with LLMOps and AI Solutions engineers to define data contracts and integration patterns between data platform and AI systems.
- Evaluate and integrate emerging data technologies that support AI workloads (vector databases, graph databases, semantic search)
Technical Leadership & Strategy (20%)
- Operate in high-ambiguity environments by defining problem statements, success metrics, and technical approach.
- Influence technology decisions and contribute to organizational data and platform strategy.
- Mentor and coach data engineers across the team, establishing technical growth paths and skill development.
- Drive adoption of engineering best practices including code review standards, testing frameworks, and CI/CD maturity.
- Represent data engineering in cross-functional planning with AI engineering, analytics, and business teams.
- Evaluate emerging technologies, frameworks, and engineering approaches for potential adoption.
- Contribute to enterprise-wide data strategy, platform roadmap, and architecture governance discussions
Operational Excellence & Governance (10%)
- Define and enforce data governance policies including classification, access control, lineage, and retention.
- Establish monitoring, alerting, and SLA frameworks for critical data pipelines and data products.
- Drive incident response processes, root cause analysis, and preventive measures for data platform issues.
- Implement cost optimization strategies for cloud data infrastructure and compute resources.
- Ensure compliance with regulatory requirements and responsible data practices
Collaboration & Stakeholder Engagement (10%)
- Partner with business stakeholders, product owners, and domain experts to translate business needs into data solutions.
- Align with AI engineering teams (LLMOps, AI Solutions) to ensure seamless data-to-AI integration.
- Communicate technical decisions, trade-offs, and roadmap to both technical and non-technical audiences.
- Facilitate knowledge transfer and documentation practices across the data engineering team
Day-to-Day Activities
A typical week in this role involves a blend of strategic architecture work, hands-on engineering, mentorship, and cross-functional collaboration:
- Designing data models and pipeline architectures for new business requirements, drawing architecture diagrams and writing technical design documents.
- Writing and reviewing complex data transformation code in Databricks (Python, Spark, SQL) for high-volume and high-complexity use cases.
- Leading technical design sessions with the team to solve ambiguous data challenges and establish patterns for reuse.
- Reviewing pull requests and providing architectural guidance and mentorship to team members.
- Meeting with AI engineering teams (LLMOps, AI Solutions) to align on data contracts, embedding pipeline requirements, and vector store integration.
- Investigating and resolving complex production issues including performance bottlenecks, data quality degradations, and pipeline failures.
- Evaluating new tools and technologies (e.g., vector databases, streaming frameworks, governance tooling) through proof-of-concepts.
- Participating in architecture review boards and contributing to enterprise data strategy discussions.
- Defining and tracking data platform KPIs (pipeline reliability, data freshness, query performance, cost efficiency).
- Conducting 1:1 mentoring sessions with junior and mid-level data engineers on technical growth.
- Collaborating with data governance and security teams on access policies, data classification, and compliance requirements.
- Writing and maintaining engineering playbooks, architectural decision records, and onboarding documentation
- Provide technical mentorship and guidance across the AI gineering organization.
- Support knowledge sharing, cross-training, and engineering excellence initiatives
Qualifikationen Technical Skills & Tools
Education
Bachelor's degree in Computer Science, Computer Engineering, Data Science, Information Systems, or equivalent. Master's degree preferred
Required Knowledge, Skills & Abilities
- 4-6+ years of progressive data engineering experience with increasing scope and complexity.
- Deep expertise in Python, SQL, Scala, and distributed data systems (Apache Spark, Delta Lake).
- Strong experience with Azure Cloud, Databricks Lakehouse platform, Unity Catalog, and Databricks Workflows.
- Experience with Docker, Kubernetes, and CI/CD pipelines for data engineering workloads.
- Advanced monitoring and logging (Azure Monitor, Log Analytics, Databricks observability).
- Proven ability to design and deliver enterprise-scale data architectures (data lake, warehouse, lakehouse).
- Experience with data architectures supporting AI/ML workloads (e.g., vector databases, embedding pipelines, unstructured data processing).
- Strong data modeling skills (dimensional, data vault, lakehouse medallion patterns).
- Experience defining and enforcing data governance, quality, and security standards.
- Track record of mentoring engineers and establishing engineering best practices.
- Excellent communication skills with ability to influence technical decisions across teams
Preferred Knowledge, Skills & Abilities
- Experience with Databricks Mosaic AI, Unity Catalog for AI governance, and Feature Store.
- Familiarity with RAG architectures, vector search (e.g., Pinecone, Weaviate, Azure AI Search), and embedding models.
- Experience with streaming data platforms (Kafka, Event Hubs, Spark Structured Streaming).
- Knowledge of LangChain or similar frameworks (understanding data requirements, not application development).
- Experience with graph databases and knowledge graph architectures.
- Infrastructure-as-code experience (Terraform, ARM templates, Bicep).
- Basic experience with Power BI or similar tools for semantic-layer-based data discovery.
- Familiarity with SAFe Agile or similar scaled agile frameworks.
- Experience operating in regulated environments with data compliance requirements
Additional Information
- Must be able to work flexible hours/work schedule.
- Travel domestically and internationally as needed.
- Work holidays and weekends when required for production support or critical deliveries.
- Enjoys collaborative work and technical mentoring with peers and junior team members.
- A self-starter who thrives in ambiguous environments and drives clarity through technical leadership.
- This role is part of MBUSA's Data Insights & AI organization and contributes to the company's data and AI strategy and operating model