Data Infrastructure Engineer (Query Engine)

zaimler

$135K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software engineering with a focus on distributed systems.
  • Proficiency in Java, Scala, Rust, or C++ programming languages.
  • Deep understanding of distributed systems architecture and query engine internals.
  • Experience with big data technologies like Apache Spark, Presto, or Trino.
  • Strong SQL skills, particularly in query optimization techniques.

Responsibilities

  • Design and develop scalable, fault-tolerant query engines.
  • Optimize performance using techniques like vectorized processing and caching.
  • Integrate with modern data lake formats like Apache Iceberg and Delta Lake.
  • Innovate solutions for concurrency, scalability, and reliability in distributed systems.
  • Contribute to open-source platforms to meet product needs.
  • Collaborate with cross-functional teams to align technical solutions with business requirements.
  • Research and implement advancements in query processing.

Benefits

  • Competitive compensation package.
  • Equity offered in the company.
  • Full medical, dental, and vision benefits.
  • 401(k) retirement plan available.
  • Flexible onsite work hours to encourage collaboration.
Full Job Description
Who You Are

  • Thrives in early-stage environments, eager to build robust systems from scratch.
  • Passionate about distributed systems and solving complex data challenges at scale.
  • Able to navigate ambiguity and adapt to changing requirements in a fast-paced startup.
  • Advocates for engineering efficiency and continuous improvement.
  • A leader who enjoys mentoring others and fostering a strong engineering culture.
  • Excited to work cross-functionally with a team that values transparency, purpose-driven innovation, and collective leadership.


What You Will Be Doing

  • Design and Develop: Build scalable, fault-tolerant query engines optimized for performance and resource efficiency.
  • Optimize Performance: Apply advanced techniques such as vectorized processing, cost-based optimization, and caching to enhance query execution.
  • Integrate Seamlessly: Develop integrations with modern data lake formats (e.g., Apache Iceberg, Delta Lake, and Hudi) and semantic layers.
  • Innovate in Distributed Systems: Architect solutions to handle concurrency, scalability, and reliability in distributed environments.
  • Leverage Open Source: Contribute to or extend platforms like Apache Spark, Presto, and Trino to meet unique product requirements.
  • Collaborate: Work with product, data science, and engineering teams to align technical solutions with business needs.
  • Stay Current: Research and implement the latest advancements in query processing and distributed systems.


Prior Experience

  • Proficiency in programming languages such as Java, Scala, Rust, or C++.
  • Deep understanding of query engine internals, distributed systems architecture, and parallel query processing.
  • Experience with modern big data technologies (e.g., Apache Spark, Presto, Trino) and data formats like Parquet, ORC, or Avro.
  • Proven ability to build and optimize scalable systems capable of processing petabyte-scale datasets.
  • Strong grasp of SQL semantics, execution plans, and query optimization techniques.


Nice to Have

  • Experience in building AI/ML infrastructure and ML production systems at scale.
  • Hands-on experience with Linux, Docker, and other containerization technologies.
  • Prior experience at an early-stage startup, developing systems and processes from scratch.


We9re a fast-moving, well-funded startup based in San Mateo, working onsite with flexible hours because the best ideas happen when smart people collaborate in person. We take ownership of our work, move with urgency while maintaining quality, and focus on delivering real results-not just effort. We offer competitive compensation, equity, full benefits (Medical, Dental, Vision, 401k), and a workspace built for collaboration, transparency, and deep technical problem-solving.

We sponsor H-1B visas and assist with immigration processes to bring the best minds together!

Similar Jobs

More Jobs at zaimler

More Information Technology Jobs

Find similar Data Infrastructure Engineer (Query Engine) jobs: