Databricks

Sr. Software Engineer - Ingestion Core team

Databricks$166K — $225K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in production code using Java, Scala, Go, C++, or Python
  • Experience in architecting and deploying large-scale distributed systems
  • Background in streaming, Spark, databases, or change data capture (CDC)
  • Expertise in optimizing systems for scale, throughput, latency, and reliability
  • Familiarity with AI tools and experience in defining evaluation metrics for development

Responsibilities

  • Build distributed infrastructure for ingesting data from various sources
  • Reduce end-to-end latency and increase the throughput of data ingestion
  • Design and optimize streaming workloads for cost and efficiency
  • Explore ML techniques to enhance streaming workload performance
  • Develop monitoring tools to improve visibility of ingestion workflows
  • Collaborate with other teams to facilitate AI use cases
  • Create strong evaluations for Agentic SDKs

Benefits

  • Comprehensive benefits and perks
  • Employee-focused programs and resources
  • Eligibility for annual performance bonus
  • Equity options for employees
  • Specific regional benefits disclosures available
Full Job Description
To enable all of this on Databricks, making data ingestion seamless is crucial. That's the mission of the Ingestion Core Team: to make the ingestion of all data-structured and unstructured-simple, reliable, and efficient. Simplifying the complex is hard, and that's where you come in. This role requires building distributed platform systems to incrementally ingest high-volume, petabyte-scale data from diverse sources-including cloud storage (SQS, ADLS, GCS), databases (Oracle, SQL Server, MySQL, Postgres), and file sources (Google Drive, SharePoint)-at high throughput and low cost. The data includes structured formats (JSON, Parquet, CSV) as well as unstructured data (text, images, docs, PPTs, and blobs), all of which land in Delta Lake with schema evolution and change data capture (CDC) capabilities.

Join us in making data ingestion effortless and be part of the team that powers the future of AI + data at Databricks!

As an engineer on the team, you will work on projects that:
  • Build distributed infrastructure to ingest data from diverse sources and support streaming ingestion, incremental processing, and replication. This isn't just about building plugin connectors.
  • Reduce end-to-end latency, increase throughput, and reduce costs from the time data appears in source systems to when it is available in Delta Lake.
  • Design and optimize streaming and distributed workloads for throughput, cost, latency, reliability, and scale.
  • Optimize streaming workloads by exploring and applying ML techniques.
  • Build monitoring and observability capabilities (customer-facing and internal) that provide visibility into ingestion workflows and the systems running them.
  • Collaborate with partner teams to enable use cases like RAG and AI agents.
  • Build Agentic SDKs with strong evaluations.


Ideal Engineer should have:
  • 5+ years of experience writing production code in one of: Java, Scala, Go, C++, or Python.
  • Experience architecting, developing, and deploying large-scale distributed and asynchronous systems.
  • Experience with distributed systems, streaming, Spark, databases, data processing, or CDC.
  • Experience building or operating systems where scale, throughput, latency, reliability, and cost are important considerations.
  • Comfortable using AI tools and defining effective evals to shorten the development loop.


Pay Range Transparency

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.

Local Pay Range

$166,000-$225,000 USD

BenefitsAt Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.

About Databricks

Databricks is a unified analytics platform that provides data engineering, collaborative data science, and machine learning capabilities. The company was founded in 2013 by the original creators of Apache Spark, a popular open-source big data processing engine. Databricks provides a cloud-based platform that allows data teams to collaborate and build data pipelines, run machine learning models, and perform advanced analytics. The company has raised over $1 billion in funding and is valued at $38 billion as of November 2021.
Learn more about Databricks
Size
2,000 employees
Industry
Founded
2013

Similar Jobs

More Jobs at Databricks

More Information Technology Jobs

Find similar Sr. Software Engineer - Ingestion Core team jobs: