Company:Qualcomm Canada ULC
Job Area:Engineering Group, Engineering Group > Machine Learning Engineering
General Summary:Today, more intelligence is moving to end devices, and mobile is becoming the pervasive AI platform. Building on the smartphone foundation and the scale of mobile, Qualcomm envisions making AI ubiquitous-expanding beyond mobile and powering other end devices, machines, vehicles, and things.
We are inventing, developing, and commercializing power-efficient on-device AI, edge cloud AI, and 5G to make this a reality.
New PositionPurpose:As a member of Qualcomm's ML Systems Team, you will:
- Maintain local server racks and on-site devices
- Make infrastructure and processes robust, reliable, and efficient
- Identify and remedy items impacting productivity of the development team.
Responsibilities:- Operate and improve the Linux self-hosted GitHub runner fleet: capacity, scheduling, storage, monitoring, recovery, access, and incident response.
- Build reliable GitHub Actions pipelines and reproducible Docker environments for builds, tests, model benchmarks, artifacts, and releases. Maintain the Git LFS-backed model zoo and its shared caches.
- Own and innovate on the Python task/workflow orchestration infrastructure (Prefect-esque) to make hardware measurements traceable, repeatable, and actionable.
- Steward performance-data ingestion and analysis, including the FastAPI/PostgreSQL-backed service and its clients; improve pytest integration, regression detection, reporting, triage, and release promotion.
- Plan the next scale step: isolate workloads, improve cache and artifact lifecycle, and evaluate elastic/cloud or batch execution where it fits scarce devices and large models.
- Turn project-specific tooling into supported, reusable platform components for other teams.
Required:Depth in:
The ideal candidate will be familiar with:- Linux infrastructure: Self-hosted GitHub Actions runners, systemd, remote filesystems (NFS), and resource monitoring.
- CI/CD and containers: GitHub Actions, reusable workflows, Docker, and release automation. Familiarity with Jenkins is a plus.
- Model and artifact management: Git LFS, shared caches, and large-model storage (ONNX models).
- Python testing: pytest, pytest-xdist, integration tests, and performance reporting (with run-to-run variation).
- AI performance tooling: ONNX, PyTorch, QAIRT SDK is a plus, Android device execution and profiling is a plus (adb).
- Performance data: REST APIs, FastAPI, and PostgreSQL, or similar libraries/frameworks.
- Future scaling: batch scheduling (ex. IBM Spectrum LSF) or cloud infrastructure for horizontally scaling automation.
Minimum Qualifications:• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 8+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
OR
Master's degree in Computer Science, Engineering, Information Systems, or related field and 7+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
OR
PhD in Computer Science, Engineering, Information Systems, or related field and 6+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.