The Role
We are seeking a Senior Software Engineer to build the applied systems that bring AI capabilities into production: the data pipelines, LLM integrations, agentic workflows, tool interfaces, and evaluation frameworks that make AI-driven products reliable enough to depend on. This is applied engineering rather than research; success is measured in shipped, maintainable systems. You may be a strong fit if you love working within a landscape that changes quickly, creating durable architectures with swappable parts, so new models and techniques are adopted on evidence.
The role encompasses fluency across cloud architecture and local model inference, evaluation design, agentic workflows and orchestration, API and tool design, and AI-assisted data enrichment. It also requires the engineering rigor to establish reliable sources of truth, detect regressions, and recognize when a deterministic solution is more appropriate than an AI-driven one.
Morningstar's hybrid work environment gives you the opportunity to collaborate in-person each week as we've found that we're at our best when we're purposely together on a regular basis. In most of our locations, our hybrid work model is four days in-office each week. A range of other benefits are also available to enhance flexibility as needs change. No matter where you are, you'll have tools and resources to engage meaningfully with your global colleagues.
This position is based in our Toronto office. We follow a hybrid policy of at least 4 days onsite.
Job Responsibilities
Integrate with hosted LLM inference (via centralized model gateway infrastructure) for extraction, classification, and agent orchestration workloads; evaluate open-weight models against hosted options for cost and performance tradeoffs.
Design and maintain adapters that sync source systems (CMS, event platforms, editorial, research libraries) into a governed dataset without duplicating source-of-record logic.
Qualifications
Hands-on production experience integrating LLMs: Skills, MCP, RAG, structured outputs, and tool/function calling.
Strong proficiency in Python across eval tooling, service-level code, and API development (FastAPI or similar), plus working proficiency in TypeScript/Node.js for application integrations.
Security and privacy judgment in AI systems: handling sensitive data appropriately, and designing against failure modes like prompt injection, data leakage through prompts, and unsafe or unattributed model output.
Experience operating production systems against latency, reliability, and cost targets, including token-cost management for inference workloads.
Evaluation literacy: you can describe an eval you built, what it caught, and how you handled canonical truth, variance, and regression cases.
Explicitly not required: formal model training or ML research credentials (e.g., pretraining, fine-tuning research). This is an applied systems role; we're looking for builders, not researchers.
Nice to have
Experience with cloud-hosted inference services (e.g., Amazon Bedrock, Azure AI/Cognitive Services) and centralized LLM gateways (e.g., LiteLLM).