About the RoleWe are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion.
The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath.
This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.
In this role, you will:- Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency.
- Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together. Tie each expansion to the workloads it can serve.
- Lead cross-stack programs connecting ingestion and processing, databases and indexes, and file/object storage. Make data ownership, schema compatibility, change-data-capture, replay and consumer-readiness contracts explicit so the full data path remains correct and usable.
- Make cost and efficiency architectural inputs. Evaluate physical versus logical footprint, index and replication amplification, redundant copies, tiering, caching and network movement against the cost of serving useful workloads.
- Drive resilience and recovery programs with explicit failure scenarios and validation. Distinguish database backup, failover and point-in-time recovery from execution/workspace save-and-restore; verify correctness, recovery time, safe resumption and isolation from live traffic.
- Coordinate lifecycle correctness across files, objects, databases and data platforms, including metadata, retention, deletion and snapshots. Incorporate privacy, access control, auditability and residency requirements into the design and consumer contracts.
- Lead adoption and major migrations through compatibility checks, representative workload testing, staged cutovers, rollback and operational handoff. Improve APIs, guardrails and self-service so new capacity and capabilities can be consumed predictably.
- Measure architecture changes through product and platform outcomes: task completion and continuity, data freshness, query/retrieval and snapshot latency, throughput, reliability and cost/efficiency. Use those results to drive durable performance and operational improvements.
You might thrive in this role if you:- Have independently owned complex production programs in data platforms, databases or storage infrastructure and can explain the architectural decisions, your contribution and the resulting impact.
- Have deep working knowledge of hyperscaler/cloud storage technologies, such as Amazon S3 or Azure Blob Storage, including their performance, placement, resiliency and cost constraints.
- Understand the full data and storage stack: product access patterns and APIs; ingestion and processing; databases, storage engines and indexes; caching, replication and data movement; and the durable-storage, CPU and network layers that support them.
- Can translate model, product and data-consumer needs into precise platform requirements, challenge assumptions and define evidence that a capability, recovery path or scaling change is ready.
- Have led delivery across product, model, data, database, storage and infrastructure teams, especially where no single team owns the end-to-end result.
- Can use workload evidence to make tradeoffs across capability, correctness, reliability, latency and cost/efficiency, and build mechanisms that improve decisions as requirements change.
- Communicate complex decisions clearly and maintain ownership through production adoption, not only a launch milestone.