Multimodal Intelligence Systems Engineer

MaxIT Consulting

$110K — $130K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 1-5 years of relevant professional experience
  • Hands-on experience deploying computer vision or vision-language systems
  • Strong understanding of visual AI for images or video
  • Experience overseeing technical systems end to end
  • Comfortable in hardware-constrained environments
  • Solid software engineering fundamentals

Responsibilities

  • Build and deploy vision-language systems for industrial workflows
  • Develop multimodal reasoning capabilities for images and video
  • Utilize detection, segmentation, and visual reasoning techniques
  • Own model orchestration and deployment in constrained environments
  • Design and maintain evaluation systems with ground-truth datasets
  • Iterate on production models based on usage and system performance

Benefits

  • Opportunity to work in an early-stage technology company
  • Engagement in real-world industrial AI applications
  • Collaborate with a team focused on cutting-edge multimodal technologies
  • Potential for impact on production-level AI systems
  • Conducive work environment in the San Francisco Bay Area
Full Job Description
Acerca del puesto Multimodal Intelligence Systems Engineer

SF Bay Area, California | Fully On-site

We are seeking an Multimodal Intelligence Systems Engineer to join an early-stage technology company building AI systems for real-world industrial environments. This role focuses on multimodal visual intelligence running in hardware-constrained environments and supporting complex physical workflows.

The Opportunity

You will own critical parts of the AI layer, building and deploying vision-language systems that reason over images and video in real customer environments. The role combines applied AI, computer vision, evaluation systems, and production engineering.

Key Responsibilities
  • Build and deploy agentic vision-language systems for real-world industrial workflows.
  • Develop multimodal reasoning capabilities across images and video.
  • Work with detection, segmentation, visual reasoning, VLMs, and related applied computer vision techniques.
  • Own model orchestration and production deployment across hardware-constrained environments.
  • Design and maintain evaluation systems including ground-truth datasets, trajectory evaluation, and regression testing.
  • Iterate on production models based on real-world usage and measurable system performance.

Required Qualifications
  • Approximately 1 to 5 years of relevant professional experience.
  • Hands-on experience shipping computer vision, multimodal, or vision-language systems to production.
  • Strong practical understanding of visual AI applied to images and/or video.
  • Experience owning technical systems end to end rather than working exclusively on research prototypes.
  • Comfort operating in real-time, edge, hardware-constrained, or connectivity-constrained environments.
  • Strong software engineering fundamentals and a high level of technical ownership.

Candidate Profile

The strongest candidates are practical builders who have shipped real systems with real users. Experience in robotics, autonomous systems, drones, wearables, real-time video, edge computing, or adjacent physical-world AI domains is highly relevant.

Work Arrangement

This position is fully in person in the San Francisco Bay Area. The initial working location is in Burlingame, with the team expected to transition to a San Francisco office.

Work Authorization

Candidates must be able to work in the United States. Certain visa sponsorship, transfer, or change-of-status scenarios may be supported for candidates already located in the United States. New petitions requiring processing from abroad are not supported at this time.

Similar Jobs

More Jobs at MaxIT Consulting

More Consumer Technology Jobs

Find similar Multimodal Intelligence Systems Engineer jobs: