AI Engineer

NetBrain Technologies

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree or higher in Computer Science, AI, Electrical Engineering, or a related field; Master's or Ph.D. preferred.
  • 3+ years in software engineering or applied AI, with 2+ years in production-grade LLM or Agent applications.
  • Experience delivering at least one LLM-powered feature end-to-end, owning its operation post-launch.
  • Deep understanding of Agent architectures and LLM behaviors, with practical experience in building multi-step workflows.
  • Proficient in diagnosing and resolving production issues, including LLM failures and retrieval-quality degradation.
  • Strong skills in Python and distributed backend engineering; capable of developing APIs and services.

Responsibilities

  • Design and construct core capabilities for an enterprise-grade Agent platform.
  • Implement LLM post-training strategies and self-learning feedback loops for continual model improvements.
  • Create LLM and Agent evaluation frameworks alongside automated regression pipelines.
  • Establish observability tools for AI system performance, including metrics and alerting mechanisms.
  • Design reliable backend services for production AI workloads, optimizing for latency and throughput.
  • Lead the technical design of system-level capabilities while collaborating across teams.
  • Prototype and benchmark emerging AI technologies to enhance system capabilities.

Benefits

  • Comprehensive compensation package including 401k and medical/dental coverage.
  • Performance-driven bonus structure available.
  • Work culture that values fast-paced iteration and impact.
  • Opportunities for professional development and growth.
  • Supportive team environment focused on innovation.
Full Job Description
What We Need

We're looking for a Senior AI Engineer to design and build production-grade agent and RAG systems that power intelligent, reliable automation across our platform. This role combines hands-on engineering with system-level thinking-owning everything from architecture and evaluation to scalability, observability, and reliability in production. The ideal candidate thrives in ambiguity, moves quickly from prototype to production, and brings a strong focus on quality, safety, and real-world impact.

What You'll Do

Agent Platform Architecture
  • Design and implement core capabilities for an enterprise-grade Agent platform, including orchestration patterns such as ReAct, Plan-and-
  • Execute, and Supervisor, as well as tool execution, context and memory management, and safety guardrails.
  • Design enterprise-grade Agent execution and governance mechanisms, including Human-in-the-Loop approval workflows, multi-tenant
  • permission isolation, policy enforcement, and secure execution controls.
  • Build reusable Agent Skills, standardized tool interfaces, and a scalable tool ecosystem deeply integrated with NetBrain platform
  • capabilities and business workflows.

LLM and Model Optimization
  • Design and implement LLM post-training strategies, including domain-specific Supervised Fine-Tuning (SFT), DPO/RLHF-based
  • preference alignment, and parameter-efficient fine-tuning techniques such as LoRA, to continuously improve model performance in the
  • network operations domain.
  • Build an Agent self-learning feedback loop that converts production execution traces, user feedback, and evaluation results into high-
  • quality datasets for continuously improving prompts, skills, models, and retrieval strategies.
  • Analyze and optimize LLM behavior across areas such as instruction following, tool calling, structured output generation, contextual
  • understanding, reasoning stability, and hallucination mitigation.

Evaluation, Reliability, and Observability
  • Build production-grade LLM and Agent evaluation frameworks and automated regression pipelines, including benchmark datasets,
  • deterministic checks, LLM-as-a-Judge, tool-call validation, retrieval-quality evaluation, and end-to-end task success metrics.
  • Establish release quality gates and hallucination-detection mechanisms for AI features to prevent significant accuracy, reliability, and
  • performance regressions from reaching production.
  • Build comprehensive AI system observability capabilities, including distributed tracing, structured logging, metrics, dashboards, and
  • alerting.
  • Rapidly diagnose and resolve production AI failures, including hallucinations, incorrect tool selection, invalid tool parameters, Agent
  • loops, retrieval-quality degradation, structured-output failures, latency regressions, and unexpected model behavior changes.

Production Engineering and Technical Execution
  • Design and implement highly reliable backend services for production AI and Agent workloads, including asynchronous and concurrent
  • processing, retries, timeouts, caching, rate limiting, and fault isolation.
  • Continuously optimize latency, throughput, token consumption, and infrastructure cost to meet platform SLA requirements and support
  • large-scale production workloads.
  • Independently diagnose and resolve complex AI system issues spanning prompts, models, RAG, tools, Agent workflows, backend
  • services, and infrastructure.
  • Lead technical design for critical modules and system-level capabilities, ensuring solutions align with platform architecture, security
  • requirements, engineering standards, and product requirements.
  • Drive technical improvements based on production data, evaluation results, and benchmarks, and collaborate closely with Engineering,
  • Product, QA, and other teams to deliver solutions into production.

Applied Research and Technical Strategy
  • Prototype, benchmark, and productionize emerging technologies such as GraphRAG, Knowledge Graphs, MCP, LLM Post-Training, and
  • Agent Self-Learning to improve grounding, multi-hop reasoning, and domain expertise.
  • Continuously evaluate Agent frameworks and supporting infrastructure, including LangChain, LangGraph, AutoGen, and LlamaIndex,
  • and provide technical recommendations for platform architecture evolution and product technology strategy.
  • Stay current with developments in LLM and Agent technologies and rapidly translate promising technologies into measurable, testable,
  • and production-ready engineering capabilities.


What You Bring
  • Bachelor's degree or higher in Computer Science, Artificial Intelligence, Electrical Engineering, or a related technical field. Master's or Ph.D.
  • preferred; equivalent practical experience will also be considered.
  • 3+ years of experience in software engineering, machine learning, or applied AI, including 2+ years building, deploying, and operating production-
  • grade LLM or Agent applications. Must have delivered at least one LLM-powered feature end-to-end and owned its ongoing operation and
  • improvement after production launch.
  • Deep understanding of Agent architectures and LLM behavioral characteristics, including instruction following, tool-calling behavior, and context
  • sensitivity, with hands-on experience building multi-step workflows involving reasoning, tool execution, state management, structured outputs,
  • validation, and error recovery.
  • Proven ability to diagnose and resolve production LLM/Agent failures, including hallucinations, incorrect tool calls, retrieval-quality degradation,
  • Agent loops, structured-output failures, latency regressions, and regressions introduced by prompt or model changes.
  • Strong Python and distributed backend engineering skills, including API and service development, asynchronous and concurrent programming,
  • retries, timeouts, caching, rate limiting, testing, logging, and cross-service performance debugging.
  • Hands-on experience designing evaluation systems for LLM applications, including dataset construction, metric definition, regression testing, and
  • release quality gates.
  • Strong understanding of security risks associated with LLM and Agent applications, including prompt injection, data leakage, unsafe tool
  • execution, permission boundaries, and uncontrolled Agent autonomy.
  • Ability to independently design, implement, debug, deploy, and operate complex production systems.

Preferred Qualifications
  • Strong experience with RAG and advanced retrieval systems, including embeddings, vector and hybrid search, reranking, chunking strategies,
  • grounding and citation mechanisms, as well as multi-hop retrieval, Knowledge Graphs, or GraphRAG.
  • Familiarity with Agent frameworks such as LangGraph, LangChain, AutoGen, and LlamaIndex, as well as standardized protocols such as MCP
  • (Model Context Protocol) for connecting Agents with external tools and systems. A strong understanding of the underlying architecture is more
  • important than expertise in any specific framework.Experience designing Agent runtime mechanisms, including Human-in-the-Loop workflows such as risk classification, approval, interruption, pause
  • /resume, as well as context management for long-running Agent workflows, including state persistence, history compression, and memory
  • systems.
  • Experience with LangSmith or similar LLM observability and evaluation platforms.
  • Experience with LLM fine-tuning, including LoRA or other parameter-efficient fine-tuning techniques.
  • Experience applying LLM technologies to networking, infrastructure, cybersecurity, observability, or other complex technical domains.
  • Fluent in both English and Chinese, with strong cross-regional communication and collaboration skills.


What We Offer

Our comprehensive compensation package is vital in how we recognize our people for the impact they make on us reaching our goals as a company.

For this role, the estimated base is $150,000 - $180,000 + Bonus. The actual salary may vary based on a range of factors, including market and individual qualifications objectively assessed during the interview process.

The range listed above is a guideline and may be modified. People Experience offers a comprehensive benefits package in addition to cash compensation that includes but is not limited to 401k and medical/dental coverage. Speak with your Recruiter for more details on our Total Rewards philosophy.

#LI-BW1

NetBrain invites all interested and qualified candidates to apply for employment opportunities.

Similar Jobs

More Information Technology Jobs

Find similar AI Engineer jobs: