ResponsibilitiesPeraton is seeking a Mid-Level Cloud AI Engineer to support the development, deployment, and operation of artificial intelligence and machine learning solutions across a multi-cloud government environment serving 70+ customer tenants and growing. The environment spans AWS, Microsoft Azure, Google Cloud Platform (GCP), and Oracle Cloud Infrastructure (OCI).
Location: Remote, but must reside and perform all work within the United States
Work Hours: This position requires working online from 8:00 AM Eastern to 5:00 PM Eastern
Day to Day Roles and Responsibilities:
AI/ML Development and Deployment
- Build, train, and deploy machine learning models using managed AI/ML services across AWS (SageMaker, Bedrock), Azure (Azure ML, Azure OpenAI Service), GCP (Vertex AI), and OCI (OCI Data Science, OCI Generative AI)
- Develop and maintain ML pipelines for data ingestion, feature engineering, model training, evaluation, and deployment
- Implement model serving infrastructure including real-time inference endpoints, batch prediction workflows, and API integration patterns
- Support the integration of large language models and generative AI capabilities into government applications with appropriate guardrails and compliance controls
Data Engineering and Processing
- Design and implement data processing workflows using cloud-native services for ETL, data lake management, and feature stores
- Work with structured and unstructured data sources to prepare training datasets, ensuring data quality, lineage, and governance requirements are met
- Optimize data pipelines for performance, cost, and reliability across cloud platforms
Monitoring, Operations, and Optimization
- Monitor deployed models for performance degradation, data drift, and bias using platform-native and third-party monitoring tools
- Troubleshoot and resolve issues across AI/ML workloads, including training failures, inference latency, and resource utilization problems
- Optimize cloud resource usage and costs for AI/ML workloads including GPU/accelerator allocation and spot/preemptible instance strategies
Collaboration and Knowledge Sharing
- Collaborate with data scientists, application developers, and infrastructure engineers to operationalize AI/ML solutions
- Document AI/ML architecture decisions, deployment procedures, and operational runbooks
- Support service delivery metrics and reporting in coordination with the Service Delivery Manager and ISR Product Owner
- Adhere to Change Management procedures for all production AI/ML deployments
Qualifications
Basic Qualifications:
- Bachelors degree and 5 years of experience or an Associates degree and 7 years of experience or a High School diploma/equivalentand 9 years of experience.
- Must be a U.S. Citizen with the ability to obtain/maintain a DHS Public Trust.
- 3 to 5 years of experience in AI/ML engineering, data engineering, or applied machine learning using cloud based technologies.
- Hands on experience with managed AI/ML services on at least two of the following cloud platforms: AWS, Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
- Proficiency in Python and experience with machine learning frameworks and libraries such as TensorFlow, PyTorch, scikit learn, or equivalent technologies.
- Experience designing, building, deploying, and maintaining ML pipelines and model serving infrastructure in production cloud environments.
- Experience with cloud based AI services, including generative AI, large language models, machine learning platforms, or related AI capabilities.
- Familiarity with responsible AI practices, model governance, data governance, and compliance requirements associated with deploying AI solutions in federal government environments.
- Strong communication, analytical, problem solving, and technical documentation skills.
Preferred Qualifications:
- DHS Public Trust or higher clearance
- Relevant cloud or AI/ML certification, such as AWS Machine Learning Specialty, Azure AI Engineer Associate, Google Professional Machine Learning Engineer, OCI AI Foundations Associate, or an equivalent certification.
- Experience working with large language models, Retrieval Augmented Generation (RAG), and the integration of generative AI services.
- Familiarity with MLOps practices, processes, and tools such as MLflow, Kubeflow, SageMaker Pipelines, Azure ML Pipelines, or equivalent technologies.
- Experience using containerization technologies, including Docker and Kubernetes, to support AI/ML workloads.
- Knowledge of data governance frameworks, policies, and tools applicable to federal data environments and data handling requirements.
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, CloudFormation, or equivalent technologies.
- Additional cloud certifications across multiple cloud service providers.
- Relevant Agile certification or demonstrated experience working in Agile development environments.
Target Salary Range$104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individuals experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.