Full Job Description
About the team
The AI/ML Platform pod builds Fanatics' internal AI platform: the paved road for every team building with LLMs, MCPs, and agents. Gateway, agent runtime, registry, retrieval, evaluation,
and cost analytics, all running on a petabyte-scale lakehouse with enterprise-grade governance. You'll join a pod that owns this platform end to end. This is a Staff-level role, and a hands-on
one. You will choose the architecture, define the standards other teams build to, and stay in the code, shipping the hardest parts yourself. The decisions you make are ones the organization
lives with for years.
Responsibilities
• Own the end-to-end architecture of the AI control plane: gateway, agent runtime,registry, retrieval, evaluation, and cost. Set technical direction, sequence the roadmap,and make the build-versus-buy calls.
• Build the enterprise LLM gateway: multi-provider routing, failover, caching, rate limiting,token optimization, and key management, with a path to self-hosted open-weightmodels where cost, latency, or data residency call for it.
• Build the Agent/MCP ecosystem on a managed runtime such as AWS BedrockAgentCore: enterprise systems exposed as governed tools over MCP, a control-plane registry for discovery, ownership, versioning, and lifecycle, and a no-code path from prototype to governed production agent, surfaced through a company-wide chat portal.
• Own the unified retrieval layer and the data platform behind it: the governed front door plus the ingestion, chunking, embedding, and indexing pipelines and the vector and
search backends that feed it, kept reproducible, observable, and quality-gated at petabyte scale.
• Build the evaluation and feedback loop: offline evals, tracing, and regression gates in CI/CD, with production failures and user edits turning into new eval cases and root-cause attribution across prompts, retrieval, tools, and models.
• Own governance and cost: RBAC, agent identity, least-privilege access to enterprise data, audit logging, use-case-level cost attribution, and the enterprise AI tool portfolio, standing up to finance, executive, and EU AI Act scrutiny.
• Raise the bar: set the engineering standards, mentor engineers, and represent the architecture and its trade-offs directly to senior leadership
Qualifications
• 8 to 12 years building production software, including 2+ years on AI/ML platform. Bachelor's or Master's in Computer Science, Engineering, or a related field.
• Staff-level impact: system design across multiple services, decisions that span teams, and platforms other engineers build on.
• Strong Python and FastAPI, plus one of Java, GoLang, Scala, or TypeScript. Hands-on with Terraform, Kubernetes, Airflow, Postgres, Grafana, and Prometheus, and familiar
enough with the AWS data science toolkit and libraries like PyTorch to partner credibly with data scientists. Chat UI or internal developer portal work is welcome.
• Hands-on with LLMs and generative AI on Bedrock or Vertex: model routing, RAG, tool calling, structured outputs, and context engineering. Agentic systems and MCP server and client integrations. Agent runtimes such as AWS Bedrock AgentCore matter here, and because they are new we look for depth rather than years. Exposure to self-hosted serving with vLLM, SGLang, or llama.cpp is welcome.
• The data engineering behind retrieval: chunking, embedding, and indexing pipelines, vector or search backends such as OpenSearch, open lakehouse table formats, and a working method for measuring retrieval quality.
• LLM evaluation and observability: eval dataset design, tracing, LLM-as-judge scoring, and regression testing, with tools such as Langfuse or Arize Phoenix.
• Security and governance for AI platforms: RBAC, OAuth2/OIDC, least privilege, cost optimization, and usage metering. Familiarity with governance systems like DataHub and semantic layers, and with administering enterprise AI tools such as ChatGPT Enterprise, Claude Enterprise, or Glean, is welcome.
• Writes clearly and holds the room with executives. Open-source contributions or published work in agentic AI or LLMOps is welcome
The salary range represents base pay only and does not include short-term or long-term incentive compensation. When determining base pay as part of a final compensation package, we consider several factors such as location, experience, qualifications, and training. For information about our benefits, please visit https://benefitsatfanatics.com/
Salary Range
$190,000-$237,500 USD