STORD

Senior Computer Vision Engineer (Egocentric), Data Foundry

STORD$120K — $160K *
Consumer Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in production computer vision or perception systems, or MS/PhD with 6+ years industry experience
  • Experience with real-world data in perception system development
  • Proven expertise in egocentric vision, robotics perception, and AI data platforms
  • Strong background in computer vision fundamentals including CNNs, transformers, and depth estimation
  • End-to-end ownership experience of perception challenges, from data to deployment
  • Deep knowledge of large-scale video and multimodal datasets
  • Proven ability to transform complex problems into scalable systems with limited resources
  • Expert-level Python skills and proficiency in software engineering; C++ skills beneficial

Responsibilities

  • Define and evolve AI data products with varying quality tiers
  • Establish data quality standards for scalable datasets
  • Design and develop complex perception systems including object detection and 3D estimation
  • Implement camera systems and synchronization for warehouse environments
  • Develop automated data pipelines for efficient dataset preparation
  • Train, optimize, and deploy computer vision models on video datasets
  • Establish methodologies for evaluating model performance and quality

Benefits

  • Access to one of the largest real-world environments for AI training data
  • Direct exposure to partnerships with robotics companies and AI labs
  • Opportunity for technical ownership in shaping foundational systems
  • Collaboration with the CTO and Co-Founder for strategic input
  • Positioning at the intersection of robotics, computer vision, and AI data
Full Job Description
Build the Vision Systems Powering the Future of Physical AI.

Stord operates the largest independent e-commerce fulfillment network in the U.S. - with 20+ fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually. We are transforming this operational infrastructure into one of the most valuable sources of training data for the next generation of physical AI.

We are building a new business line at the intersection of robotics, computer vision, and AI data - and we are looking for an experienced Computer Vision Engineer to help build the technical foundation from the ground up.

This is a hands-on builder role for a technical leader who can design, prototype, and productionize perception systems that transform real-world environments into high-quality AI training data.

Why This Role:

This is a rare opportunity to build the technical foundation of a new AI business from the ground up - combining real-world operational infrastructure with cutting-edge computer vision and robotics.

You will have:
  • A structural advantage no startup can easily replicate - access to one of the largest real-world environments for collecting physical AI training data.
  • Direct exposure to the fastest-growing AI market - partnering with robotics companies, AI labs, and teams building the future of intelligent systems.
  • True technical ownership - the opportunity to define architecture, build foundational systems, and shape the future of Embodied AI data.
  • Executive partnership - working closely with Stord's CTO and Co-Founder to define strategy, accelerate execution, and remove barriers.

What You Will Own:

You will own the early computer vision and egocentric perception stack - including data capture systems, vision pipelines, model development, and the infrastructure required to deliver high-quality datasets at scale.

Working closely with a small, highly technical team, you will help define the architecture, build the systems, and establish the technical standards for Stord's Embodied AI data platform.

Build the Data Product & Capture Platform
  • Define and evolve Stord's Embodied AI data products across quality tiers - from RGB egocentric video to depth-enhanced and multimodal datasets with hand pose, body pose, and rich annotations.
  • Determine the right technical investments based on customer requirements and the needs of emerging robotics and AI models.
  • Establish data quality standards and evaluation frameworks to ensure datasets meet production-level requirements.


Build the Perception Stack
  • Design and develop perception systems including:
    • Object detection, tracking, and segmentation
    • Depth estimation and 3D reconstruction
    • 6DoF pose estimation
    • Multi-view 3D hand and body pose estimation
    • Egocentric and fixed-camera perception systems
  • Build robust solutions designed for complex, real-world environments - not just benchmark datasets.


Own the Hardware + Vision Integration
  • Design and deploy camera systems and perception rigs across warehouse environments.
  • Own camera calibration, multi-camera synchronization, epipolar geometry, and 3D reconstruction workflows.
  • Develop solutions for deriving accurate spatial understanding from multimodal sensor inputs and video data.


Build Automated Labeling & Data Pipelines
  • Develop VLM-assisted and automated annotation workflows with human-in-the-loop quality systems.
  • Integrate labeling tools and processes that improve scalability while maintaining dataset accuracy.
  • Build pipelines that transform raw video into production-ready training datasets.


Train, Optimize, and Deploy Models
  • Design, fine-tune, evaluate, and optimize computer vision and multimodal models on large-scale video datasets.
  • Build reproducible training and deployment workflows that move beyond experimentation and into production.
  • Establish evaluation methodologies that measure model performance, reliability, and quality.


What You'll Need:
  • 8+ years of experience building and shipping production computer vision or perception systems (or an MS/PhD in Computer Vision, Machine Learning, Robotics, or a related field with 6+ years of hands-on industry experience).
  • Experience building perception systems that operate on real-world, imperfect data - beyond academic benchmarks.
  • Demonstrated experience developing or scaling egocentric vision, robotics perception, autonomous systems, or AI data platforms.
  • Deep expertise in computer vision fundamentals and modern tooling, including:
    • CNNs and vision transformers
    • Object detection, segmentation, and tracking
    • Depth estimation and 3D vision
    • 2D/3D pose estimation
    • Camera calibration and geometric computer vision
  • Experience owning complex perception problems end-to-end - from data collection and model development through evaluation, optimization, and deployment.
  • Strong understanding of large-scale video and multimodal datasets, including data quality measurement and evaluation methodologies.
  • Experience taking ambiguous, 01 technical challenges and turning them into scalable systems with limited resources.
  • Ability to influence technical direction, establish engineering standards, and mentor other engineers.
  • Expert-level Python skills and strong software engineering fundamentals; C++ experience where performance requirements demand it.

About STORD

STORD is a cloud-based warehousing and distribution network that provides modern logistics infrastructure for businesses of all sizes. The company's platform connects a network of warehouses across the United States and provides businesses with real-time visibility, control, and optimization of their inventory and supply chain operations. STORD's mission is to empower businesses with the technology and infrastructure they need to compete in today's fast-paced, global economy.
Learn more about STORD
Size
100 employees
Industry

Similar Jobs

More Jobs at STORD

More Consumer Technology Jobs

Find similar Senior Computer Vision Engineer (Egocentric), Data Foundry jobs: