Req ID: 390502
We are currently seeking a Infrastructure Architect to join our team in Bay Area, Charlotte, Dallas, Phoenix, NY, North Carolina (US-NC), United States (US).
Job Description: We are seeking an experienced AI Platform Architect to lead the technical architecture, strategy, and evolution of enterprise AI platforms. This role will provide architectural leadership for large-scale AI and cloud platform initiatives, partnering with engineering, infrastructure, security, product, and business stakeholders to deliver scalable, resilient, and secure AI capabilities.
The successful candidate will demonstrate deep expertise in cloud-native architectures, AI/ML platforms, distributed systems, and platform engineering, along with the ability to influence technical roadmaps and lead cross-functional initiatives.
Key Responsibilities - Architecture Leadership
- Lead architecture and technical direction for enterprise AI platform capabilities, including:
- Enterprise Generative AI platforms
- Agentic AI platforms and agent runtime environments
- Model serving and inference infrastructure
- Prompt engineering, evaluation, and testing frameworks
- AI governance, risk management, and guardrails
- AI observability, monitoring, and operations
- Multi-cloud AI platform strategy and architecture
- Platform Architecture & Design
- Design, evaluate, and guide architecture across AI and cloud technologies such as:
- Red Hat OpenShift AI (RHOAI)
- Google Cloud Vertex AI
- Gemini models
- Azure AI Foundry
- Amazon Bedrock
- Anthropic Claude
- OpenAI services and models
- Cloud-native platform scalability and resiliency solutions
- Strategic Initiatives
- Lead and collaborate on initiatives involving:
- Capacity planning and performance optimization
- GPU infrastructure and platform strategy
- Large-scale NVIDIA-based AI infrastructure architectures
- Multi-region and multi-cloud resiliency
- Disaster recovery planning and cloud DR strategies
- Active-active platform architectures
- High availability and business continuity solutions
Required Qualifications - Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; or equivalent combination of education and relevant experience.
- Minimum of 10 years of experience in software engineering, infrastructure engineering, systems architecture, or related technical disciplines.
- Minimum of 5 years of experience designing and implementing large-scale cloud-native platforms.
- Experience designing, deploying, and supporting highly available, mission-critical production systems.
Experience with: - Kubernetes and/or OpenShift
- Distributed systems architectures
- Public cloud platforms and cloud-native architectures
- AI/ML platforms and infrastructure
- API and integration platforms
- Data platforms and data-intensive applications
Knowledge of: - Generative AI systems and architectures
- Retrieval-Augmented Generation (RAG)
- Agentic AI frameworks and platforms
- Model serving and inference architectures
- MLOps practices and tooling
- Demonstrated ability to lead technical initiatives across multiple teams and stakeholders.
- Ability to work from or relocate to one of the following approved locations: San Francisco Bay Area, CA; Charlotte, NC; Dallas, TX; Phoenix, AZ; or New York, NY.
Preferred Qualifications - Experience with one or more major cloud providers, including AWS, Azure, or Google Cloud Platform.
- Experience architecting GPU-accelerated AI infrastructure.
- Experience implementing AI governance, security, risk management, and compliance controls.
- Experience building or supporting multi-region, highly resilient enterprise platforms.
- Relevant industry certifications in cloud, AI/ML, Kubernetes, or architecture disciplines.