Deep hands-on experience with LLMs, including deployment and debugging.
Ability to thrive in a fast-paced, changing environment.
Experience building or optimizing inference systems with a strong grasp of latency, throughput, and TCO.
Skill in explaining complex ML concepts to non-technical stakeholders.
Commercial mindset with a focus on revenue expansion and strategic growth.
Responsibilities
Lead hands-on demos of Kimchi's inference products.
Set up Kimchi Studio workspaces and configure integrations for customers.
Design and implement governance policies for budget control and model routing.
Assist teams in integrating Kimchi's serverless inference API.
Translate customer feedback into actionable product insights.
Benefits
Flexible, remote-first work environment.
Collaboration with a global team of cloud experts.
Equity options available.
Fast-paced workflow with quick feedback cycles.
10% of work time dedicated to personal projects or self-improvement.
Learning budget for professional development, including conferences.
Team-building budget for company events.
Equipment budget for necessary tools.
Extra days off for work-life balance.
Full Job Description
About this role:
You'll be the founding customer-facing engineer for our flagship Kimchi product within our Customer Success organization. Kimchi executes multi-model LLM workflows and autonomous agents with embedded governance and cost control. It intelligently routes across frontier and open-source models while enabling teams to establish strict budget caps, enforce model access policies, configure roles, and audit TCO via deep analytical tracking. Through Kimchi Studio, developers gain desktop access to agentic workflows, helping organizations accelerate adoption without sacrificing operational governance or spend visibility.
This is a hands-on, individual contributor role where you will live at the intersection of inference optimization, autonomous coding agents, multi-model routing and LLM governance. You will have the unique opportunity to architect and execute the growth strategy that scales Kimchi adoption and creates new revenue expansion opportunities for our existing customers. Requirements:
Deep hands-on experience with LLMs: you've deployed them, fine-tuned them, evaluated them, or debugged their failure modes. You care about why model A outperforms model B for a specific task.
You know what it's like to operate in a fast-moving environment that's prone to change.
You've built or optimized inference systems out of curiosity. You know the difference between latency, throughput, and total cost of ownership. You ballpark costs and performance implications in your head.
You can explain complex inference concepts to non-ML engineers and help them see why it matters to their business.
You can think commercially. You understand revenue expansion, can identify expansion opportunities within customer organizations, and can design strategies to capture them. You're not just a technical person, you're someone who combines technical excellence with entrepreneurial thinking about growth.
Responsibilities:
Lead hands-on demos of our flagship inference products - Kimchi Coding, Studio and Teleport.
Set up Kimchi Studio workspaces within customer environments. Help teams configure worktrees, integrations (Slack, repos, docs), and skills so agents understand their conventions.
Design and drive implementation of governance policies for customers: spending limits per team/developer, cost attribution by project, model routing preferences. Make sure the hard budget caps actually work for their org.
Help teams integrate Kimchi's serverless inference API. Optimize model selection and routing for their specific workloads (coding, reasoning, planning tasks).
Translate customer feedback into product input. Highlight where the agent fails, where Studio UX gets in the way, where governance policy doesn't match their needs.
What's in it for you?
Enjoy a flexible, remote-first global environment.
Collaborate with a global team of cloud experts and innovators, passionate about pushing the boundaries of Kubernetes technology.
Equity options.
Get quick feedback with a fast-paced workflow. Most feature projects are completed in 1 to 4 weeks.
Spend 10% of your work time on personal projects or self-improvement.
Learning budget for professional and personal development - including access to international conferences and courses that elevate your skills.
Team-building budget and company events to connect with your colleagues.
Equipment budget to ensure you have everything you need.
Extra days off to help maintain a healthy work-life balance.
Hiring process
Screening call with Recruiter
Hiring Manager interview
1-2 additional interviews based on the role
Culture Check interview with an executive
#LI-Remote
About CAST AI
CAST is a technology corporation headquartered in New York City and in France, near Paris. The firm markets software intelligence, with a technology based on semantic analysis of software source code and components. It provides information and size measurement technology and expertise. CAST offers software, hosting and consulting services in support software analysis and measurement. The company was founded in 1990 in Paris, France, by Vincent Delaroche.
Its code quality metrics are used in application development service-level agreements, and the firm offers consultation on issues having to do with development quality and security:
CAST was founded in 1990 in Paris by Vincent Delaroche. In 1996 it shipped its first software product based on semantic analysis of Transact-SQL. CAST Application Intelligence Platform, was first launched in 2004, initially introducing software quality measurement. In 2017, CAST Highlight is launched as a SaaS product scanning portfolio of software to provide metrics on health, cloud migration capabilities, and Open-source license risks. Early 2019, based on the same analysis technology, the firm launched CAST Imaging, a product representing graphically source-code components of a software.
In 2012, the firm announced support for the Object Management Group Automated Function Point Standard, one way of measuring application development productivity.
The firm's leadership includes Bill Curtis, who developed the Capability Maturity Model at the Software Engineering Institute in the early 1990s and then the Consortium for IT Software Quality.