Job Type
Full-time
Description
About the Role:Reporting to the Director of Data & AI Engineering, the AI/ML Engineer is a mid-level, hands-on technical role responsible for building and maintaining the data pipelines, AI models, and intelligent features that power the GLOBO platform. This role spans the full lifecycle of AI development-from cleaning and preparing data, to building and evaluating models, to shipping production features that directly improve operational efficiency and customer experience.
The AI/ML Engineer works across GLOBO's modern data stack (Fivetran, dbt, Snowflake) and AI infrastructure (AWS Bedrock, LLMs, agentic frameworks) to deliver reliable, well-tested solutions. This person is equally comfortable wrangling messy data and prompt-engineering an LLM, and takes pride in writing clean, tested code that other engineers can build on.
Data Engineering & Pipeline Development: - Build and maintain reliable ingestion pipelines usingSnowflake Openflow, Python, and Snowflake, including API, PostgreSQL, and CDC-based integrations.
- Develop incremental synchronization, cursor/state management, retry logic, schema-drift handling, soft-delete propagation, and source-to-target reconciliation.
- Transform raw source data through staging, intermediate, and core models into trusted datasets for analytics, reporting, and machine-learning workloads.
- Apply data-quality checks for freshness, completeness, uniqueness, referential integrity, valid relationships, and business-rule compliance.
- Maintain source definitions, model documentation, lineage, metadata, and data contracts.
- Collaborate with data owners to ensure PII/PHI classification, masking, retention, and deletion requirements are implemented throughout the pipeline.
Model Development, Testing & Evaluation: - Implement monitoring and alerting for ingestion failures, pipeline freshness, schema changes, data-quality failures, transformation errors, model drift, and inference degradation.
- Establish automated regression testing for dbt models, features, evaluation datasets, prompts, and model outputs.
- Validate that sensitive data is appropriately masked, redacted, access-controlled, and excluded from unauthorized model training or data-sharing workflows.
- Build safeguards for PII/PHI in recorded-call, transcript, and AI/ML processing pipelines, including verification that redaction and deletion workflows complete successfully.
- Ensure AI/ML outputs are traceable to their source data, model or prompt version, feature set, and evaluation results.
- Define recovery procedures, data-quality escalation paths, and operational runbooks for critical pipelines and models.
- Support human review and approval for model outputs that may affect customers, interpreters, employees, financial activity, or service quality.
Feature Development & Integration: - Collaborate with Product and Engineering to ship AI-powered features into the GLOBO platform.
- Build and deploy LLM integrations (AWS Bedrock, Anthropic Claude) and agentic workflows (CrewAI, LangChain).
- Write production-quality code with proper tests, documentation, and error handling.
Reliability & Safety: - Implement guardrails, monitoring, and alerting for AI services in production.
- Ensure AI outputs are consistent and trustworthy.
- Contribute to evaluation datasets, prompt versioning, and regression testing for deployed models.
Performance & Cost Optimization: - Monitor and optimizeSnowflake compute, storage, query performance, dbt execution, Openflow runtime usage, and model-inference costs.
- Design efficient incremental models, CDC pipelines, materializations, clustering strategies, and warehouse/task schedules.
- Compare and optimize ingestion costs as Globo transitions from Fivetran to Snowflake Openflow.
- Reduce unnecessary full refreshes, duplicate processing, excessive data movement, and inefficient feature recomputation.
- Optimize model selection, prompt size, token usage, batching, caching, inference frequency, and routing between model providers.
- Measure model performance against operational cost, latency, throughput, and data-freshness requirements.
- Establish practical service-level targets for critical datasets, transformations, batch jobs, and model-serving workflows.
Requirements
Required Minimum Education and Experience :- Bachelor's Degree in Computer Science, Data Science, Information Systems, or related field.
- 2+ years of experience in data engineering, software development, or ML engineering.
- Experience with the below tech stack is required:
- Python (advanced proficiency)
- SQL (advanced proficiency)
- LLM Integration (AWS Bedrock, Anthropic Claude, or OpenAI API)
- dbt (data transformation and testing)
- Snowflake (or similar cloud data warehouse)
- AWS Lambda / Serverless architecture
- Experience with the below tech stack is preferred:
- Fivetran (or similar ELT/ingestion tooling)
- Agentic Frameworks (CrewAI, LangChain, or similar)
- Airflow (or similar workflow orchestration)
- Vector Databases (Pinecone, PGVector, or OpenSearch)
- AWS ECS/EKS
- CDK and CloudFormation for automated deployments
- Ruby on Rails (ability to read/debug core platform code)
- Redis
- PostgreSQL
- React
- Familiarity with model evaluation techniques, prompt engineering, and AI safety best practices.
- Experience with Google Docs and Apple/Mac Operating System preferred
Additional Preferred Requirements :- Ability to work independently in a decentralized environment without the reliance on direct authority
- Highest level of personal and professional integrity and ethics
- Broad understanding of current and emerging technology practices
- High level of initiative, accountability, and follow-through
- Value strong teamwork and collaboration skills
- Demonstrated problem-solving and decision-making skills
- Ability to manage multiple initiatives and projects and prioritize needs
- Strong sense of service and passion for the company and business
- Authorized to legally work for any employer in the United States
- Willingness to submit to any requested background checks
- Fluent in English