You will ship LLM-powered product features end-to-end. That means designing retrieval and tool-calling flows, writing the services that run them, building evals and guardrails, and watching cost, latency, and quality in production. You'll partner with the ML/infra teammates on embeddings, indexing, and model hosting, and with the product teammates on user experience and outcomes. We move fast, and we care about reliability in a safety-critical domain.
This role is for engineers who set technical direction across services and teams (8+ years experience, with a track record of owning systems or defining standards others build against). You'll lead design for ambiguous, cross-team problems like provider routing or shared retrieval infra.
What you'll doBuild user-facing LLM features- Design and implement retrieval-augmented generation and tool-calling flows using frameworks like LangChain or equivalent primitives, where simpler is better.
- Deliver robust JSON and schema-bound outputs with validation, retries, and fallbacks.
- Add function calling to integrate with internal tools, search, routing, and data services.
Own the service layer- Ship APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff.
- Add caching, request shaping, prompt templates, and context packing to control latency and cost.
- Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints as needed.
Retrieval and data prep- Collaborate with infrastructure teammates to develop chunking, embeddings, and indexing capabilities for documents, time series, and multimedia.
- Choose and tune vector backends such as OpenSearch, pgvector, or Pinecone.
- Keep knowledge bases fresh with data syncs from S3, Aurora, DynamoDB, and external sources.
Evaluation and quality- Create offline evals and golden sets for prompts, retrievers, and tools.
- Stand up online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request.
- Run A/B tests and prompt/version rollouts with guardrails and canaries.
Safety, privacy, and compliance- Implement content and policy checks, PII detection and redaction, access controls, and auditing.
- Design human-in-the-loop paths for sensitive actions.
- Handle aviation data with care and follow internal security standards.
Operate what you build- Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation.
- Debug tricky failures across retrieval, prompts, tools, and providers.
What will make you successful- Shipped LLM apps: You've put LLM features in front of users and improved them with data.
- Strong builder: Comfortable writing production code, tests, and docs. You keep things simple and observable.
- RAG and tools depth: You understand embeddings, chunking, vector search tradeoffs, and function calling.
- Quality mindset: You design evals, define success metrics, and iterate based on evidence.
- Cost and latency aware: You track p95, hit SLAs, and reduce cost without hurting quality.
- Clear communicator: You explain tradeoffs and align partners across product, infra, and security.
- Technical leadership: You've set direction or standards that other engineers or teams built against, not just shipped your own code.
Nice to have- Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.
- Prompt versioning, guardrails, and provider routing in production.
- Multimodal work with time series or video.
- Familiarity with GPU inference, Triton, or TensorRT-LLM.
- Aviation or other safety-critical domain exposure.
- DevOps basics for CI/CD, IaC, and secure secrets handling.
Example problems you might tackle in month one- Transform an internal knowledge base into a low-latency RAG service, complete with explicit schemas and evaluations.
- Add tool-calling to automate a repetitive cockpit or ops workflow with guardrails and audit trails.
- Reduce the cost per request through improved chunking, caching, and prompt refactoring, while maintaining task success rates.
Work LocationThis is a hybrid role based in San Carlos, CA, with 3+ days per week onsite and the option to work remotely on remaining days.
Perks & Benefits (Full-Time Employees)- Healthcare: 100%* of employee medical premiums covered; 25% for dependents
- Time Off: 3 weeks PTO plus 13+ paid company holidays
- 401(k): Offered (no current employer match, but we are committed to enhancing this benefit in the future)