7+ years of software engineering experience, with 2+ years in LLM-based agents.
Expertise in LLM application architecture, including model selection and orchestration.
Proficient in Python, with knowledge of function calling and evaluation tooling.
Experience in building benchmarks for task completion and cost efficiency.
Hands-on technical leadership experience in a small team setting.
Background in applying AI/LLM tools to engineering data or software, particularly in mechanical engineering or CAD/CAE/PLM domains.
Familiarity with desktop automation and enterprise workstation constraints.
Responsibilities
Own and improve agent task success metrics and performance baselines.
Develop evaluation infrastructure based on validated user stories and workflows.
Manage token budgets and track workflow costs.
Collaborate with researchers to validate user workflows through interviews.
Translate user stories into testable evaluations, prioritizing by value and feasibility.
Make architectural decisions regarding tool calling and context management.
Lead a small team by reviewing designs, unblocking teammates, and contributing code.
Align agent behavior with real-world use through collaboration with product and customers.
Benefits
Comprehensive health insurance coverage.
Flexible work hours and potential for remote work.
Professional development opportunities and training.
Collaborative and innovative work environment.
Full Job Description
About the Role
Lead the development of agent intelligence that helps mechanical engineers complete complex, multi-step workflows across desktop engineering software. As a hands-on technical lead on a small AI engineering team, you will shape the agent architecture, connect user research with product development, and improve the reliability and cost efficiency of real-world workflows. What You'll Do
Own agent task success metrics, establish performance baselines, and systematically improve completion rates.
Build rigorous, reproducible evaluation infrastructure grounded in validated user stories and real engineering workflows.
Set per-task token budgets and track the cost of completed workflows.
Work with researchers and engineering domain experts to map, interview about, and validate user workflows.
Translate validated user stories into testable evaluations and prioritize workflow coverage by customer value and technical feasibility.
Make architecture decisions across tool calling, state management, error recovery, model routing, and context management.
Set technical direction for a small team, review designs and code, unblock teammates, and contribute production code.
Collaborate with product, integrations, and customers to align agent behavior with real-world use.
What We're Looking For
At least 7 years of software engineering experience, including 2 or more years building and shipping LLM-based agents that take real-world actions.
Deep experience with LLM application architecture, including model selection, context management, retrieval, tool calling, and orchestration.
Strong Python skills and familiarity with function calling, tool APIs, tracing or observability, and evaluation tooling.
Experience building benchmarks for task completion, cost efficiency, and failure analysis.
Hands-on technical leadership and code review experience on a small engineering team.
Experience applying AI or LLM tooling to proprietary engineering data or desktop engineering software, with background in mechanical engineering, CAD, CAE, PLM, or a related domain.
Familiarity with desktop automation and enterprise workstation constraints is valuable.
Compensation & Benefits
Salary range: $160,000 to $250,000 USD annually. Visa sponsorship is not available. Location
On-site in San Francisco, California, United States.