Full Job Description
Core AI engineering (must-have)
LLM application design. Prompt engineering, tool/function calling, structured output, and agent patterns, plus judgment about when an LLM is the wrong tool.
Retrieval-augmented generation (RAG). Chunking strategies, embeddings, vector stores, retrieval tuning, and grounding answers with citations.
Evaluation. Building test sets and LLM-as-judge or golden-answer evals, and measuring hallucination, accuracy and latency. This is the biggest gap between people who've done demos and people who've shipped.
Responsible AI and guardrails. PII and PHI handling, which matters a lot in insurance and disability claims. Also prompt-injection defense,
Languages. Strong Python, plus either TypeScript/Node or Java. Java helps here because of the Spring Boot / JBoss estate.
CI/CD for AI assets. Versioning prompts, bot definitions, contact flows and knowledge base content the same way as code. Promoting Lex bots and Connect flows Dev QA UAT Prod through pipelines. Running evals as a pipeline quality gate.
Amazon Connect. Contact flows, routing profiles and queues, Lambda integrations, Contact Lens (transcripts and sentiment), Amazon Q in Connect (agent assist), CTR/contact data streaming.
Amazon Lex V2. Intents, slots, utterance design, confidence thresholds, fallback handling, versions and aliases, and multi-language bots.
Amazon Bedrock.
Workflow orchestration. Step Functions, EventBridge, Lambda, SQS/SNS, and knowing when to use Bedrock Flows or Agents versus plain Step Functions.
Platform fundamentals. IAM (least privilege, cross-account roles; Lex and Connect access is tightly controlled here), VPC basics, CloudWatch and X-Ray, S3, DynamoDB, KMS, and cost awareness for token and Connect usage.
Infrastructure as code. Terraform preferred (or CDK/CloudFormation). Connect and Lex are notoriously hard to manage as code
Azure DevOps and software engineering
- Azure Repos. Git branching strategies (feature/release/hotfix), PR workflow, branch policies, code review discipline.
- Azure Pipelines. YAML multi-stage pipelines, templates, variable groups, service connections to AWS (OIDC or role assumption), environments with approvals and gates, and artifact feeds.
- Application Health Monitoring. Use of Data Dog and or CloudWatch to monitor an application's health, provide warning and alert messages for proactive management. Manage incidents include defining incident playbooks.