Staff AI Engineer, Applied AI - Smart Vision

Wepay

$150K — $180K *
Consumer Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS in Computer Science or related field; MS/PhD in ML or CV preferred.
  • 8+ years in building production ML/AI systems with end-to-end model ownership.
  • Strong expertise in computer vision: detection, classification, segmentation, tracking.
  • 3+ years experience with LLMs or VLMs in production environments.
  • Proficient in Python and PyTorch; solid knowledge of cloud and distributed systems (AWS, Docker).
  • Demonstrated ability in evaluation and maintaining data quality in ML projects.
  • Experience with embeddings and video retrieval at scale.

Responsibilities

  • Build and fine-tune computer vision models for various detection and video understanding tasks.
  • Own the pipeline for video-understanding using vision-language models, focusing on quality tuning.
  • Adapt and optimize models to specific domain requirements, ensuring efficiency and effectiveness.
  • Manage the data loop, including curation and evaluation to enhance model accuracy.
  • Oversee the embedding and retrieval system for video search, optimizing query handling.
  • Design agentic experiences on the vision stack, ensuring effective communication and reasoning.
  • Optimize production inference for cost efficiency and speed while maintaining performance.
  • Collaborate with cross-functional teams to develop and successfully roll out AI features.

Benefits

  • Opportunity to influence and shape cutting-edge AI projects.
  • Collaborative and innovative work environment with diverse teams.
  • Commitment to professional growth and support for development in key areas.
Full Job Description
About the role

As a Staff AI Engineer for Applied AI, you'll be a technical owner of the models behind Arlo's smart features - from computer vision detectors running on-camera to vision-language models that describe what happened, to the retrieval and agent layers that let customers ask questions about their video. You'll pick the right approach for each problem (train, fine-tune, prompt, or retrieve), prove it with solid evals, and take it all the way to production at consumer scale. This is a hands-on applied role: you ship models, not papers.

What you'll do
  • Build, train, and fine-tune computer vision models for detection, classification, tracking, re-identification, and video understanding, and improve them against real-world customer footage - night, weather, motion blur, odd camera angles, edge compute limits.
  • Own our video-understanding pipeline built on vision-language models: frame selection and temporal context, prompt and output-schema design, grounding and hallucination control, multi-event reasoning, and quality tuning for captioning and scene description.
  • Adapt models to our domain: SFT, LoRA/QLoRA, preference tuning, distillation into small deployable models, and knowing when a 200M-parameter specialist beats a frontier model.
  • Own the data and evaluation loop - dataset curation, labeling strategy, hard-negative and failure mining, active learning, benchmark suites, and offline/online metrics that reliably predict customer-perceived quality.
  • Own the embedding and retrieval stack behind video search: multimodal/video embeddings, vector index design and tuning, hybrid search and re-ranking, and natural-language queries over a user's library.
  • Build agentic experiences on top of the vision stack: tool/function calling, multi-step reasoning over event history, RAG and memory, guardrails, and tracing/observability for agent runs.
  • Keep production inference fast and economical - serving stack choice and tuning (vLLM/TensorRT-LLM/Triton), quantization, batching, GPU utilization, and routing between hosted foundation models and self-hosted open models against clear cost and latency targets.
  • Partner with product, data, backend, and firmware/edge teams to turn ambiguous product ideas into shipped AI features, with safe rollout (canary, A/B, feature flags) and real production telemetry.
  • Raise the engineering bar - architecture and design reviews, MLOps practices, documentation, and mentoring senior and mid-level engineers.


What we're looking for
  • BS in Computer Science or a related technical field with 8+ years of experience; MS/PhD in ML, CV, or a related field preferred (or equivalent practical experience).
  • 8+ years building production ML/AI systems, with a track record of owning models end to end - problem framing, data, training, evaluation, deployment, iteration.
  • Strong computer vision depth: detection, classification, segmentation, tracking, video understanding; you've trained and shipped CV models in a real product.
  • 3+ years working with LLMs or VLMs in production - multimodal modeling, prompt and context design, fine-tuning, and evaluation.
  • Strong Python + PyTorch; solid distributed systems and cloud fundamentals (AWS, Docker).
  • Rigorous about evaluation and data quality - you build the benchmark before you build the model.
  • Experience with embeddings and vector search at scale; multimodal or video retrieval strongly preferred.
  • Experience shipping LLM agents with tool use, and pragmatic judgment about when not to use an agent.
  • Working knowledge of inference optimization (serving stacks, quantization, batching, GPU performance) and cost/latency tradeoffs at scale.
  • Demonstrated technical leadership without formal authority: influencing roadmaps, mentoring engineers, driving cross-team decisions.
  • Bias to ship, comfort with ambiguity, strong written communication.


Nice to have
  • Edge/on-device inference (quantization-aware training, pruning, NPU/DSP toolchains) and streaming inference.
  • Video processing at scale (decoding, frame sampling, FFmpeg-class tooling).
  • Audio or sensor-fusion models complementing video; privacy-preserving or on-device personalization.
  • Open-source contributions to CV, serving, retrieval, or agent frameworks; consumer IoT or camera/security domain experience.


We're committed to inclusivity and selecting the strongest candidate-no matter their background. Even if you don't meet every listed qualification, we encourage you to apply. We're happy to support growth in areas essential to the role. Interested in learning more about our workplace? Visit and follow our LinkedIn, and Glassdoor pages to read employee insights and get updates of what it's like to be part of Arlo.

Similar Jobs

More Jobs at Wepay

More Consumer Technology Jobs

Find similar Staff AI Engineer, Applied AI - Smart Vision jobs: