Build the Vision Systems Powering the Future of Physical AI.
Stord operates the largest independent e-commerce fulfillment network in the U.S. - with 20+ fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually. We are transforming this operational infrastructure into one of the most valuable sources of training data for the next generation of physical AI.
We are building a new business line at the intersection of robotics, computer vision, and AI data - and we are looking for an experienced Computer Vision Engineer to help build the technical foundation from the ground up.
This is a hands-on builder role for a technical leader who can design, prototype, and productionize perception systems that transform real-world environments into high-quality AI training data.
Why This Role:This is a rare opportunity to build the technical foundation of a new AI business from the ground up - combining real-world operational infrastructure with cutting-edge computer vision and robotics.
You will have:
- A structural advantage no startup can easily replicate - access to one of the largest real-world environments for collecting physical AI training data.
- Direct exposure to the fastest-growing AI market - partnering with robotics companies, AI labs, and teams building the future of intelligent systems.
- True technical ownership - the opportunity to define architecture, build foundational systems, and shape the future of Embodied AI data.
- Executive partnership - working closely with Stord's CTO and Co-Founder to define strategy, accelerate execution, and remove barriers.
What You Will Own:You will own the early computer vision and egocentric perception stack - including data capture systems, vision pipelines, model development, and the infrastructure required to deliver high-quality datasets at scale.
Working closely with a small, highly technical team, you will help define the architecture, build the systems, and establish the technical standards for Stord's Embodied AI data platform.
Build the Data Product & Capture Platform- Define and evolve Stord's Embodied AI data products across quality tiers - from RGB egocentric video to depth-enhanced and multimodal datasets with hand pose, body pose, and rich annotations.
- Determine the right technical investments based on customer requirements and the needs of emerging robotics and AI models.
- Establish data quality standards and evaluation frameworks to ensure datasets meet production-level requirements.
Build the Perception Stack- Design and develop perception systems including:
- Object detection, tracking, and segmentation
- Depth estimation and 3D reconstruction
- 6DoF pose estimation
- Multi-view 3D hand and body pose estimation
- Egocentric and fixed-camera perception systems
- Build robust solutions designed for complex, real-world environments - not just benchmark datasets.
Own the Hardware + Vision Integration- Design and deploy camera systems and perception rigs across warehouse environments.
- Own camera calibration, multi-camera synchronization, epipolar geometry, and 3D reconstruction workflows.
- Develop solutions for deriving accurate spatial understanding from multimodal sensor inputs and video data.
Build Automated Labeling & Data Pipelines- Develop VLM-assisted and automated annotation workflows with human-in-the-loop quality systems.
- Integrate labeling tools and processes that improve scalability while maintaining dataset accuracy.
- Build pipelines that transform raw video into production-ready training datasets.
Train, Optimize, and Deploy Models- Design, fine-tune, evaluate, and optimize computer vision and multimodal models on large-scale video datasets.
- Build reproducible training and deployment workflows that move beyond experimentation and into production.
- Establish evaluation methodologies that measure model performance, reliability, and quality.
What You'll Need:- 8+ years of experience building and shipping production computer vision or perception systems (or an MS/PhD in Computer Vision, Machine Learning, Robotics, or a related field with 6+ years of hands-on industry experience).
- Experience building perception systems that operate on real-world, imperfect data - beyond academic benchmarks.
- Demonstrated experience developing or scaling egocentric vision, robotics perception, autonomous systems, or AI data platforms.
- Deep expertise in computer vision fundamentals and modern tooling, including:
- CNNs and vision transformers
- Object detection, segmentation, and tracking
- Depth estimation and 3D vision
- 2D/3D pose estimation
- Camera calibration and geometric computer vision
- Experience owning complex perception problems end-to-end - from data collection and model development through evaluation, optimization, and deployment.
- Strong understanding of large-scale video and multimodal datasets, including data quality measurement and evaluation methodologies.
- Experience taking ambiguous, 01 technical challenges and turning them into scalable systems with limited resources.
- Ability to influence technical direction, establish engineering standards, and mentor other engineers.
- Expert-level Python skills and strong software engineering fundamentals; C++ experience where performance requirements demand it.