Responsibilities- Build and operate the data and AI platform behind TeamViewer's agentic products, including ingestion, transformation, embeddings, indexing, retrieval, warehouses, lakes, and vector stores.
- Own production infrastructure for model workloads, covering deployment, versioning, routing, caching, rate limiting, cost governance, and provider management.
- Build observability for AI systems, including agent traces, tool calls, quality signals, latency, token cost, failure-mode analysis, and AI-specific alerting.
- Implement CI/CD for AI systems so prompts, tool definitions, retrieval configuration, evaluation suites, and model versions are tested, deployed, monitored, and rolled back safely.
- Run evaluation infrastructure for AI engineers, including dataset management, harness execution, regression tracking, and release reporting inside the delivery pipeline.
- Define and enforce data quality, governance, security, privacy, access control, encryption, and sensitive-data handling across structured and unstructured sources.
- Design experimentation platforms that allow the team to test AI and retrieval changes safely against real traffic.
- Collaborate with AI engineers, software engineers, product, security, and platform teams to turn requirements into reliable production systems.
Requirements- 8+ years of industry experience with strong Python expertise, solid SQL knowledge, sound software engineering fundamentals, and a proven track record of building production-grade data pipelines and platform services.
- Hands-on experience operating model-based applications in production, including retrieval pipelines, embeddings, vector databases, retrieval optimization, deployment, and observability.
- Strong understanding of MLOps and LLMOps practices, including model and prompt versioning, evaluation frameworks, tracing, regression tracking, and AI-specific quality monitoring.
- Proven ability to optimize AI workloads across providers and architectures by balancing cost, latency, reliability, and quality through effective caching and deployment strategies.
- Experience with CI/CD, automated testing, and major cloud platforms, with Azure or GCP preferred.
- Solid understanding of data governance, security, privacy, and GDPR requirements within AI-powered systems and workflows.
- Regular use of AI coding agents, combined with a critical review mindset and accountability for the correctness, security, and maintainability of delivered software.
- Practical experience with agentic development environments and extension models, including custom tools, MCP servers, repository-level instruction files, and sub-agents.
- Deep understanding of common AI failure modes, including hallucinations, context degradation, prompt injection, non-determinism, and silent regressions, along with effective mitigation approaches.
- Strong problem-solving skills, the ability to work independently, experience debugging complex systems, and a pragmatic approach to engineering trade-offs, coupled with clear communication and fluency in English.
What we offer- Competitive compensation and bonuses
- Flexible PTO and paid holidays
- 401(k) with employer matching
- Comprehensive Health insurance package including 100% employer-paid medical coverage
- Up to 12 weeks of Parental Leave
- Basic Life Insurance, Short-Term & Long-Term Disability, 100% employer-paid
- Quarterly teambuilding events, leadership luncheons, and companywide "All Hands" meetings
- Open door policy and business casual dress code
Work location for this position is Austin, TX.
Department Research & Development Locations Austin Remote status Hybrid Employment type Full-time Type of Job Non Student