Must Have Technical/Functional Skills
Strong cloud engineering and AI/ML platform engineering experience.
Kubernetes, containers, Infrastructure as Code, CI/CD, and API technologies.
Python and software engineering experience.
Experience with ML/LLM deployment and GenAI frameworks.
Knowledge of vector databases, embedding models, model APIs, and AI orchestration.
Experience with monitoring, telemetry, security, and production support.
Programming: Advanced proficiency in Python; additional experience in Go, C++, or Rust is often preferred.
Infrastructure & Cloud: Hands-on expertise with Kubernetes, Docker, and cloud platforms like AWS (SageMaker, Bedrock) or Azure (AI Foundry, OpenAI).
AI Frameworks: Familiarity with orchestration and LLM tooling such as LangChain, LangGraph, Ray, or Kubeflow.
Infrastructure-as-Code (IaC): Knowledge of automation tools like Terraform and Ansible.
Roles & Responsibilities
• Build and maintain enterprise AI/ML and GenAI platform capabilities.
• Implement model endpoints, gateways, orchestration layers, and AI services.
• Build infrastructure supporting LLMs, SLMs, RAG, embeddings, vector stores, and AI agents.
• Develop automated deployment and CI/CD pipelines for AI applications and models.
• Implement LLMOps/MLOps capabilities covering deployment, monitoring, evaluation, versioning, and observability.
• Implement model-routing and inference-management capabilities.
• Support token, compute, latency, and inference-cost optimization.
• Integrate AI platforms with enterprise identity, security, logging, monitoring, networking, and API-management services.
• Automate infrastructure provisioning and platform configuration.
• Establish production reliability, scalability, resilience, and operational standards.
• Partner with AI architects, application engineers, data engineers, and security teams.
• Platform Architecture: Design and deploy scalable, cloud-native or bare-metal infrastructure (using Kubernetes, OpenShift, or AWS/Azure) to support large language models (LLMs) and machine learning workloads.
• Model Operations (MLOps/LLMOps): Optimize model serving, inference performance, GPU utilization, and automated pipelines for fine-tuning and deploying AI models.
• Agentic & GenAI Integration: Build and maintain shared services like AI gateways, Retrieval-Augmented Generation (RAG) frameworks, and autonomous agent orchestration platforms.
• Governance & Security: Enforce enterprise data protection, compliance, auditability, and cybersecurity standards across all AI tooling.
• Developer Enablement: Create self-service platforms, APIs, and monitoring/observability tools so internal software and data science teams can safely adopt AI capabilities.
TCS Employee Benefits Summary:
Discretionary Annual Incentive.
Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
Family Support: Maternal & Parental Leaves.
Insurance Options: Auto & Home Insurance, Identity Theft Protection.
Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimburseme nt.
Time Off: Vacation, Time Off, Sick Leave & Holidays.
Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.
Salary Range $110,000 - $130,000 a year