About the team: The Knowledge Platform team is at the heart of our company's data and AI strategy. We are building the foundational infrastructure that empowers the entire company to leverage data, AI, and ML into business impact. We accomplish this by creating platforms that handle the entire data and AI/ML lifecycle, simplifying the interface while providing proper safety and governance. This includes managing our data infrastructure (Databricks, Spark, Kafka, etc.), the technology to serve that data to our users (RAG, MCP, etc.), and the platform to host and govern these AI/ML models.
In 2026, our team's overarching mission is to evolve our full data ecosystem-encompassing both platform and models-into a fully AI agent-ready infrastructure; we will empower customers to engage directly with the data platform to extract actionable value through capabilities like analytics and natural language querying, while also upgrading the platform to deliver robust, real-time performance for instant, data-driven decision-making.
What you'll doIn this role, you will lead the architecture and construction of core infrastructure components, taking full ownership of technical decision-making and custom development from foundational storage to high-availability service layers. You will perform deep-dives into systems internals and performance tuning-focusing on state management, checkpointing, and exactly-once semantics-to ensure millisecond-level accuracy and reliability across 30+ global regions. A significant portion of your impact will involve developing the "Knowledge Platform" to power AI-driven products through scalable vector indexing, agentic workflows, and real-time data streaming.
Beyond the initial build, you will own the full SDLC for high-performance distributed systems that process petabytes of data at a global scale. As a champion of engineering rigor, you will advocate for elite software practices and contribute to shared tooling, SDKs, and automation frameworks that enhance productivity across the entire organization. Finally, you will serve as a technical beacon and mentor for mid-level engineers, driving high-quality project delivery through meticulous design reviews and hands-on technical leadership.
This is a hybrid role based in San Francisco.
Who you are:- Experience: Have 5+ years of experience building and operating large-scale distributed systems or infrastructure platforms.
- First Principles Thinking: Possess a deep understanding of computer science fundamentals, including distributed systems, memory management, and networking protocols.
- The "Builder" Stack: Are proficient in Java, Kotlin, or Go. You should have hands-on experience (or the desire to deep-dive) into the internals of Kafka, Flink, Spark, Kubernetes, or OLAP engines.
- Ownership Mindset: Have a proven track record of taking 0-1 ownership of complex technical challenges, from initial design to production stability.
- Curiosity & Impact: Are intrinsically motivated to explore emerging tech in AI/ML infrastructure and real-time systems to create tangible business impact.
Preferred Qualifications:- Hands-on design experience in crafting data processing patterns for a modern Lakehouse architecture.
- Contribute to the design and development of standard framework modules, high-performance services, and client libraries for big data using tools like GCP, Databricks, BigQuery, DataProc, Kafka, Kubernetes, Spark, DataFlow, Google Cloud Storage, and Airflow.
- Excellent written and verbal communication skills tailored for diverse audiences (leadership, users, company-wide).
- Ability to rapidly evaluate various technologies and conduct proof of concepts to drive architecture design.
- Experience thriving in a complex environment.