STORD

Senior Computer Vision Engineer (Egocentric), Data Foundry

STORD$120K — $160K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience building production computer vision or perception systems, or an MS/PhD with 6+ years of industry experience
  • Experience building perception systems using real-world data, beyond academic benchmarks
  • Demonstrated expertise in egocentric vision, robotics perception, or AI data platforms
  • Deep knowledge of computer vision fundamentals and modern tooling, including CNNs and vision transformers
  • Proficiency in complex perception challenges from data collection to model deployment
  • Strong understanding of large-scale video and multimodal datasets, including data quality evaluation
  • Expert-level Python programming skills with C++ experience for performance-critical applications.

Responsibilities

  • Own the early computer vision and egocentric perception stack
  • Define and evolve Embodied AI data products across quality tiers
  • Establish standards and frameworks for data quality evaluation
  • Design and develop perception systems for object detection and tracking
  • Integrate camera systems across warehouse environments
  • Develop automated annotation workflows with human-in-the-loop quality systems
  • Train and optimize computer vision and multimodal models on large datasets.

Benefits

  • Opportunity to shape the technical foundation of a new AI business
  • Access to a large real-world environment for data collection
  • Direct collaboration with the CTO and Co-Founder
  • Ownership of architecture and foundational systems
  • Exposure to cutting-edge robotics and AI partnerships.
Full Job Description
Build the Vision Systems Powering the Future of Physical AI. Stord operates the largest independent e-commerce fulfillment network in the U.S. - with 20+ fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually. We are transforming this operational infrastructure into one of the most valuable sources of training data for the next generation of physical AI. We are building a new business line at the intersection of robotics, computer vision, and AI data - and we are looking for an experienced Computer Vision Engineer to help build the technical foundation from the ground up. This is a hands-on builder role for a technical leader who can design, prototype, and productionize perception systems that transform real-world environments into high-quality AI training data. **Why This Role:** This is a rare opportunity to build the technical foundation of a new AI business from the ground up - combining real-world operational infrastructure with cutting-edge computer vision and robotics. You will have: - **A structural advantage no startup can easily replicate** - access to one of the largest real-world environments for collecting physical AI training data. - **Direct exposure to the fastest-growing AI market** - partnering with robotics companies, AI labs, and teams building the future of intelligent systems. - **True technical ownership** - the opportunity to define architecture, build foundational systems, and shape the future of Embodied AI data. - **Executive partnership** - working closely with Stord's CTO and Co-Founder to define strategy, accelerate execution, and remove barriers. **What You Will Own:** You will own the early computer vision and egocentric perception stack - including data capture systems, vision pipelines, model development, and the infrastructure required to deliver high-quality datasets at scale. Working closely with a small, highly technical team, you will help define the architecture, build the systems, and establish the technical standards for Stord's Embodied AI data platform. **Build the Data Product & Capture Platform** - Define and evolve Stord's Embodied AI data products across quality tiers - from RGB egocentric video to depth-enhanced and multimodal datasets with hand pose, body pose, and rich annotations. - Determine the right technical investments based on customer requirements and the needs of emerging robotics and AI models. - Establish data quality standards and evaluation frameworks to ensure datasets meet production-level requirements. **Build the Perception Stack** - Design and develop perception systems including: - Object detection, tracking, and segmentation - Depth estimation and 3D reconstruction - 6DoF pose estimation - Multi-view 3D hand and body pose estimation - Egocentric and fixed-camera perception systems - Build robust solutions designed for complex, real-world environments - not just benchmark datasets. **Own the Hardware + Vision Integration** - Design and deploy camera systems and perception rigs across warehouse environments. - Own camera calibration, multi-camera synchronization, epipolar geometry, and 3D reconstruction workflows. - Develop solutions for deriving accurate spatial understanding from multimodal sensor inputs and video data. **Build Automated Labeling & Data Pipelines** - Develop VLM-assisted and automated annotation workflows with human-in-the-loop quality systems. - Integrate labeling tools and processes that improve scalability while maintaining dataset accuracy. - Build pipelines that transform raw video into production-ready training datasets. **Train, Optimize, and Deploy Models** - Design, fine-tune, evaluate, and optimize computer vision and multimodal models on large-scale video datasets. - Build reproducible training and deployment workflows that move beyond experimentation and into production. - Establish evaluation methodologies that measure model performance, reliability, and quality. **What You'll Need:** - 8+ years of experience building and shipping production computer vision or perception systems (or an MS/PhD in Computer Vision, Machine Learning, Robotics, or a related field with 6+ years of hands-on industry experience). - Experience building perception systems that operate on real-world, imperfect data - beyond academic benchmarks. - Demonstrated experience developing or scaling egocentric vision, robotics perception, autonomous systems, or AI data platforms. - Deep expertise in computer vision fundamentals and modern tooling, including: - CNNs and vision transformers - Object detection, segmentation, and tracking - Depth estimation and 3D vision - 2D/3D pose estimation - Camera calibration and geometric computer vision - Experience owning complex perception problems end-to-end - from data collection and model development through evaluation, optimization, and deployment. - Strong understanding of large-scale video and multimodal datasets, including data quality measurement and evaluation methodologies. - Experience taking ambiguous, 0→1 technical challenges and turning them into scalable systems with limited resources. - Ability to influence technical direction, establish engineering standards, and mentor other engineers. - Expert-level Python skills and strong software engineering fundamentals; C++ experience where performance requirements demand it.

About STORD

STORD is a cloud-based warehousing and distribution network that provides modern logistics infrastructure for businesses of all sizes. The company's platform connects a network of warehouses across the United States and provides businesses with real-time visibility, control, and optimization of their inventory and supply chain operations. STORD's mission is to empower businesses with the technology and infrastructure they need to compete in today's fast-paced, global economy.
Learn more about STORD
Size
100 employees
Industry

Similar Jobs

More Jobs at STORD

More Information Technology Jobs

Find similar Senior Computer Vision Engineer (Egocentric), Data Foundry jobs: