What We NeedWe're looking for a Senior AI Engineer to design and build production-grade agent and RAG systems that power intelligent, reliable automation across our platform. This role combines hands-on engineering with system-level thinking-owning everything from architecture and evaluation to scalability, observability, and reliability in production. The ideal candidate thrives in ambiguity, moves quickly from prototype to production, and brings a strong focus on quality, safety, and real-world impact.
What You'll DoAgent Platform Architecture- Design and implement core capabilities for an enterprise-grade Agent platform, including orchestration patterns such as ReAct, Plan-and-
- Execute, and Supervisor, as well as tool execution, context and memory management, and safety guardrails.
- Design enterprise-grade Agent execution and governance mechanisms, including Human-in-the-Loop approval workflows, multi-tenant
- permission isolation, policy enforcement, and secure execution controls.
- Build reusable Agent Skills, standardized tool interfaces, and a scalable tool ecosystem deeply integrated with NetBrain platform
- capabilities and business workflows.
LLM and Model Optimization- Design and implement LLM post-training strategies, including domain-specific Supervised Fine-Tuning (SFT), DPO/RLHF-based
- preference alignment, and parameter-efficient fine-tuning techniques such as LoRA, to continuously improve model performance in the
- network operations domain.
- Build an Agent self-learning feedback loop that converts production execution traces, user feedback, and evaluation results into high-
- quality datasets for continuously improving prompts, skills, models, and retrieval strategies.
- Analyze and optimize LLM behavior across areas such as instruction following, tool calling, structured output generation, contextual
- understanding, reasoning stability, and hallucination mitigation.
Evaluation, Reliability, and Observability- Build production-grade LLM and Agent evaluation frameworks and automated regression pipelines, including benchmark datasets,
- deterministic checks, LLM-as-a-Judge, tool-call validation, retrieval-quality evaluation, and end-to-end task success metrics.
- Establish release quality gates and hallucination-detection mechanisms for AI features to prevent significant accuracy, reliability, and
- performance regressions from reaching production.
- Build comprehensive AI system observability capabilities, including distributed tracing, structured logging, metrics, dashboards, and
- alerting.
- Rapidly diagnose and resolve production AI failures, including hallucinations, incorrect tool selection, invalid tool parameters, Agent
- loops, retrieval-quality degradation, structured-output failures, latency regressions, and unexpected model behavior changes.
Production Engineering and Technical Execution- Design and implement highly reliable backend services for production AI and Agent workloads, including asynchronous and concurrent
- processing, retries, timeouts, caching, rate limiting, and fault isolation.
- Continuously optimize latency, throughput, token consumption, and infrastructure cost to meet platform SLA requirements and support
- large-scale production workloads.
- Independently diagnose and resolve complex AI system issues spanning prompts, models, RAG, tools, Agent workflows, backend
- services, and infrastructure.
- Lead technical design for critical modules and system-level capabilities, ensuring solutions align with platform architecture, security
- requirements, engineering standards, and product requirements.
- Drive technical improvements based on production data, evaluation results, and benchmarks, and collaborate closely with Engineering,
- Product, QA, and other teams to deliver solutions into production.
Applied Research and Technical Strategy- Prototype, benchmark, and productionize emerging technologies such as GraphRAG, Knowledge Graphs, MCP, LLM Post-Training, and
- Agent Self-Learning to improve grounding, multi-hop reasoning, and domain expertise.
- Continuously evaluate Agent frameworks and supporting infrastructure, including LangChain, LangGraph, AutoGen, and LlamaIndex,
- and provide technical recommendations for platform architecture evolution and product technology strategy.
- Stay current with developments in LLM and Agent technologies and rapidly translate promising technologies into measurable, testable,
- and production-ready engineering capabilities.
What You Bring- Bachelor's degree or higher in Computer Science, Artificial Intelligence, Electrical Engineering, or a related technical field. Master's or Ph.D.
- preferred; equivalent practical experience will also be considered.
- 3+ years of experience in software engineering, machine learning, or applied AI, including 2+ years building, deploying, and operating production-
- grade LLM or Agent applications. Must have delivered at least one LLM-powered feature end-to-end and owned its ongoing operation and
- improvement after production launch.
- Deep understanding of Agent architectures and LLM behavioral characteristics, including instruction following, tool-calling behavior, and context
- sensitivity, with hands-on experience building multi-step workflows involving reasoning, tool execution, state management, structured outputs,
- validation, and error recovery.
- Proven ability to diagnose and resolve production LLM/Agent failures, including hallucinations, incorrect tool calls, retrieval-quality degradation,
- Agent loops, structured-output failures, latency regressions, and regressions introduced by prompt or model changes.
- Strong Python and distributed backend engineering skills, including API and service development, asynchronous and concurrent programming,
- retries, timeouts, caching, rate limiting, testing, logging, and cross-service performance debugging.
- Hands-on experience designing evaluation systems for LLM applications, including dataset construction, metric definition, regression testing, and
- release quality gates.
- Strong understanding of security risks associated with LLM and Agent applications, including prompt injection, data leakage, unsafe tool
- execution, permission boundaries, and uncontrolled Agent autonomy.
- Ability to independently design, implement, debug, deploy, and operate complex production systems.
Preferred Qualifications- Strong experience with RAG and advanced retrieval systems, including embeddings, vector and hybrid search, reranking, chunking strategies,
- grounding and citation mechanisms, as well as multi-hop retrieval, Knowledge Graphs, or GraphRAG.
- Familiarity with Agent frameworks such as LangGraph, LangChain, AutoGen, and LlamaIndex, as well as standardized protocols such as MCP
- (Model Context Protocol) for connecting Agents with external tools and systems. A strong understanding of the underlying architecture is more
- important than expertise in any specific framework.Experience designing Agent runtime mechanisms, including Human-in-the-Loop workflows such as risk classification, approval, interruption, pause
- /resume, as well as context management for long-running Agent workflows, including state persistence, history compression, and memory
- systems.
- Experience with LangSmith or similar LLM observability and evaluation platforms.
- Experience with LLM fine-tuning, including LoRA or other parameter-efficient fine-tuning techniques.
- Experience applying LLM technologies to networking, infrastructure, cybersecurity, observability, or other complex technical domains.
- Fluent in both English and Chinese, with strong cross-regional communication and collaboration skills.
What We OfferOur comprehensive compensation package is vital in how we recognize our people for the impact they make on us reaching our goals as a company.
For this role, the estimated base is CAD $130,000 - CAD $165,000 + Bonus. The actual salary may vary based on a range of factors, including market and individual qualifications objectively assessed during the interview process.
The range listed above is a guideline and may be modified. People Experience offers a comprehensive benefits package in addition to cash compensation that includes but is not limited to RRSP and medical/dental coverage. Speak with your Recruiter for more details on our Total Rewards philosophy.
#LI-BW1