Full Job Description
AI Solution/Principal Engineer
We are looking for an AI Solution/Principal Engineer to join a team building a sovereign, multi-tenant agentic AI platform along with the applied products on top of it. Everything ships under strict data residency constraints and works in Arabic and English. Our squads are small and senior, and we need someone who owns how the organization knows its AI systems work - evaluation across assistants, retrieval pipelines, agent workflows, voice and document systems, and the infrastructure that turns "it seems better" into evidence someone can act on. This is not a QA role with AI added. Test automation and performance testing are in scope, but the center of gravity is evaluating non-deterministic systems, where the same input produces different outputs, "correct" is often fuzzy, and failure modes such as hallucination, grounding failure, drift, and prompt injection are ones no assertion library catches.
Responsibilities
Set evaluation strategy: what gets measured, at which layer, with what methodology, and how results feed product decisions
Build the shared platform, including evaluation harnesses, golden-set management, dataset versioning, automated grading, regression detection, and reporting leadership can read
Make grading trustworthy through judge model selection, rubric design, calibration against human labels, and knowing when automated grading cannot be trusted
Evaluate retrieval and agents for grounding and citation correctness, tool-use validation, multi-step reasoning and failure recovery
Develop Arabic golden sets and judges calibrated for Arabic rather than assumed to transfer from English
Run online evaluation and drift detection
Conduct prompt injection, jailbreak and data-leakage red-teaming
Wire evaluation and test automation together as quality gates in CI/CD
Requirements
5+ years of experience building evaluation or test infrastructure others depend on, including harnesses, shared libraries and frameworks
At least 1 year of relevant leadership experience
Expertise in AI evaluation covering output quality, retrieval and grounding, and regression detection on non-deterministic behavior
Proficiency in LLM-as-judge methodology: rubric design, calibration against human labels, and understanding its failure modes
Advanced proficiency in Python, at the depth needed to build shared libraries, with intermediate competency in TypeScript or Java
Knowledge of statistics for non-deterministic systems, including sampling, confidence intervals, inter-rater agreement and significance
Skills in CI/CD framework design across API, web and data surfaces, with test selection, parallelization and flake management
Background in golden sets and dataset versioning
English proficiency at an Upper-Intermediate level (B2) or higher
Nice to have
Familiarity with ragas, DeepEval, Promptfoo, Braintrust, LangSmith or Langfuse
Expertise in Arabic evaluation, including golden sets, dialect coverage, right-to-left validation and judge calibration
Background in adversarial and security testing, including red-teaming
Capability to perform voice and conversational evaluation, and performance testing against AI services
Experience delivering in a regulated or government environment