Function
Cloud & Data Engineering
Job descriptionMeet Our TeamJoin a forward-thinking Engineering and AI Platform team focused on building the next generation of enterprise AI solutions. Our team is pioneering agentic AI ecosystems powered by AWS Bedrock, AgentCore, MCP servers, and modern cloud-native technologies.
As a Senior AWS AgentCore Platform Engineer, you'll work alongside Cloud Architects, AI Engineers, Platform Engineers, and Security specialists to establish scalable, secure, and observable AI platforms. You'll play a critical role in defining the operational foundation that enables enterprise teams to deploy AI agents with confidence, governance, and efficiency.
This is an exciting opportunity to shape enterprise AI infrastructure, drive innovation in LLMOps, and influence platform standards across multiple business units.
What You'll Be DoingAI Platform Observability & Reliability
- Design and implement enterprise-grade observability solutions for AI agent ecosystems built on AWS Bedrock, AgentCore, and MCP servers.
- Assess and optimize CloudWatch, X-Ray, Bedrock logging, and AgentCore tracing capabilities against agentic workflow requirements.
- Conduct gap analyses and implement observability solutions using Dynatrace and other monitoring platforms.
- Develop distributed tracing frameworks for AI workloads, including:
- LLM decision paths
- Tool invocations
- Sub-agent interactions
- MCP server communications
- Build structured logging frameworks to support troubleshooting, governance, and performance optimization.
- Design post-deployment validation pipelines for AI agents and MCP servers, including deployment health monitoring and registration verification.
Cost Governance & Optimization
- Architect cost visibility and governance frameworks across AI workloads.
- Extend cloud tagging strategies to include agent runtimes, vector databases, MCP services, and Bedrock token consumption.
- Develop cost allocation models to provide spending transparency by team, department, and application.
- Build dashboards and reporting solutions for AI platform cost tracking and forecasting.
- Configure AWS Budgets, automated alerts, anomaly detection, and optimization recommendations.
- Deliver automated cost reporting through Microsoft Teams and email channels.
Monitoring & Incident Management
- Define enterprise monitoring standards and alerting frameworks across AI platform services.
- Create and manage P1-P4 alerting strategies covering:
- Deployment failures
- Runtime exceptions
- Tool invocation errors
- MCP connectivity issues
- Integrate monitoring and notification workflows with Microsoft Teams and email.
- Develop operational runbooks and self-service documentation within Confluence.
- Evaluate AWS-native and third-party monitoring solutions and recommend target-state architectures.
Security & Platform Governance
- Assess IAM architectures and multi-team access models for enterprise-scale AI environments.
- Design Attribute-Based Access Control (ABAC) frameworks to support secure multi-team isolation.
- Evaluate Cedar policy engine capabilities within AgentCore for fine-grained authorization models.
- Develop reusable Terraform modules to enforce governance, security, and compliance standards.
- Identify scalability risks and implement secure platform design patterns for enterprise AI adoption.
Platform Engineering & Automation
- Build and maintain Infrastructure-as-Code solutions using Terraform.
- Design and enhance CI/CD pipelines supporting AI platform deployments.
- Collaborate with engineering, security, architecture, and business stakeholders in Agile environments.
- Drive platform standardization, automation, and operational excellence initiatives.
What You'll Bring to the TeamRequired Qualifications
- 8+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud services including:
- IAM
- CloudWatch
- AWS Lambda
- AWS Bedrock
- Cloud-native monitoring and governance services
- Hands-on experience implementing observability and distributed tracing solutions using tools such as:
- Dynatrace
- Jaeger
- Honeycomb
- OpenTelemetry
- Experience designing and managing Infrastructure-as-Code using Terraform.
- Strong background building and maintaining CI/CD pipelines in enterprise environments.
- Experience working in Agile teams utilizing Microsoft Teams, Confluence, and modern collaboration tools.
Preferred Qualifications
- Experience supporting AI, Generative AI, or LLM-based platforms.
- Familiarity with AgentCore, LangChain, LangFuse, LiteLLM, MCP servers, or similar AI orchestration frameworks.
- Understanding of LLM lifecycle management, prompt execution flows, token consumption tracking, and AI workload optimization.
- Knowledge of cloud cost management, FinOps practices, and governance frameworks.
- Experience designing enterprise-scale security architectures using ABAC and policy-based authorization models.
- Strong analytical and problem-solving skills with the ability to translate complex technical challenges into scalable platform solutions.
Success Factors
- Passion for emerging AI technologies and cloud-native engineering.
- Ability to balance reliability, security, performance, and cost optimization.
- Strong communication skills with the ability to influence technical and business stakeholders.
- Proven ability to lead platform modernization initiatives and establish engineering best practices.
Location: Reading, PA (Hybrid – 2-3 days onsite per week)