Credit Acceptance Corporation

Staff Machine Learning Engineer, Platform (MLOPS)

Credit Acceptance Corporation • $154K — $226K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, Statistics, or a relevant technical field with 5-7 years of experience.
  • 5+ years of experience building and operating production ML or AI systems, with ownership of training or inference pipelines, model serving, or monitoring.
  • Demonstrated experience managing a production ML or AI service through its full lifecycle.
  • Strong proficiency in Python and SQL, with a focus on production-quality engineering practices.
  • Hands-on experience with cloud ML platforms, preferably AWS and Databricks.

Responsibilities

  • Own the end-to-end deployment path for ML and GenAI models, including pipelines and model registry.
  • Monitor and ensure runtime health for models in production, including incident response and quality detection.
  • Operate the agent runtime layer, ensuring secure and governed access to production agents.
  • Evaluate agents in production through online scoring and behavioral monitoring.
  • Build and maintain observability and evaluation infrastructure for other teams' dependencies.

Benefits

  • 401(K) match
  • Adoption assistance
  • Parental leave
  • Tuition reimbursement
  • Comprehensive medical, dental, and vision coverage
  • Unique nonstandard benefits that enhance workplace culture.
Full Job Description
In this role you will own and operate the platform that every ML and AI model at Credit Acceptance runs on. Pipelines, serving, registry, evaluation infrastructure, monitoring and cost. Your job is to make the path from a working model to a reliable production capability short, repeatable and observable, so that product and science teams ship without rebuilding infrastructure each time. This is an operations and infrastructure role, not a modeling role. Success is measured by what stays up, what deploys safely, what is measurable in production, and what the platform costs to run. **Outcomes and Activities:** - **This position will work from home; occasional planned travel to an assigned Southfield, Michigan office location may be required. However, this position is permitted to work at a Southfield, Michigan office location if requested by the team member** - Own the deployment path for ML and GenAI models end to end: training and inference pipelines, model registry and versioning, serving endpoints, and controlled promotion across development, QA and production. - Own runtime health for models in production: monitoring, alerting, drift and quality-regression detection, latency and throughput objectives, capacity and autoscaling behavior, and incident response through to root cause and a closed corrective action. - Operate the agent runtime layer. Route production agents through the enterprise AI Gateway and MCP Gateway rather than direct model and tool access, migrate existing agents onto that governed path, and keep tool surfaces scoped, versioned and least-privilege as they change. - Own evaluation of agents in production, not just before release. Online scoring and behavioral monitoring, quality-regression and drift detection against pinned baselines, sampling and judge pipelines, and the release-gate mechanics that stop a regression from shipping. - Build and maintain the observability and evaluation substrate other teams depend on: trace and telemetry capture including multi-turn and multi-step agent traces, logging standards, evaluation pipeline plumbing, and the data contracts underneath them. - Own platform unit economics. Measure and manage cost per inference, per document and per interaction, and produce the platform and infrastructure cost analysis that informs build-versus-buy and hosting decisions. - Make the paved road real. Deliver reusable pipeline templates, deployment patterns, reference implementations and internal tooling so product teams adopt the standard path because it is faster, not because it is mandated. - Partner with Cloud Engineering, Data Engineering, Security and SRE so the ML platform sits inside enterprise governance, identity and observability rather than beside it. - Respond to AI-specific production incidents and drive them to a closed corrective action: prompt injection attempts, rogue-agent cost spikes, data-classification exposure through a tool call, delegation abuse between agents, and model endpoint failures. - Maintain the architecture documentation and system diagrams for the ML platform, and keep them accurate enough to be used in design review. - Mentor engineers and interns on production ML practice, and raise the operating standard through design and code review rather than through rework. **Competencies:** The following items detail how you will be successful in this role. - **Customer Empathy:** Customer Empathy is the ability to understand the perspectives, pain points, and experiences of customers. It involves actively putting oneself in the customer's shoes, comprehending their needs and challenges, and using that understanding to provide a better, more customer-centric experience. - **Engineering Excellence:** Engineering Excellence is about bringing great craftsmanship and thought leadership to deliver an outstanding product that delights customers and solves for the business. This involves the pursuit and achievement of high standards, best practices, innovation, and superior solutions. - **One Team:** A One Team mindset refers to a collaborative approach across the organization, where individuals work together seamlessly, without boundaries, as a single, cohesive team. Shared goals, open communication and mutual support create a sense of collective purpose. This enables teams to navigate challenges and pursue shared objectives more effectively. - **Owner's Mindset:** Owner's Mindset involves adopting a set of behaviors that reflect a sense of responsibility, accountability, strategic thinking, and a proactive approach to managing your domain. As an owner, you understand the business and your domain(s) deeply and solve for the right outcome for the domain(s) and the business. **Requirements** - Bachelor's degree in Computer Science, Engineering, Statistics or a relevant technical field with at least 7 years of relevant experience, or a Master's degree in one of those fields with at least 5 years of relevant experience. - 5+ years building and operating production ML or AI systems, with direct ownership of at least two of: training or inference pipelines, model serving infrastructure, model registry and versioning, or production monitoring and alerting. - Demonstrated ownership of a production ML or AI service through its full operating life: deployed it, monitored it, was paged for it, diagnosed a failure, and shipped the fix. - Strong Python and SQL, with production-quality engineering practice: version control, testing, code review, and CI/CD applied to ML workloads rather than only to application code. - Hands-on experience with a cloud ML platform in production. AWS and Databricks strongly preferred, including model serving, job orchestration, and a model registry or experiment tracking system such as MLflow. - Working knowledge of how LLM and GenAI workloads differ operationally from traditional ML: token cost and latency behavior, caching and batching, non-deterministic output. - Experience running LLM or agent applications behind a gateway or proxy layer, and able to build best practices for model routing and fallback, credential and key management, rate limiting and budget enforcement. - Working knowledge of tool-calling architecture for agents, including the Model Context Protocol: what an MCP server is, how tools are scoped and authorized, and why tool access is brokered through a gateway rather than granted directly. - Experience with containerization and infrastructure as code. - Ability to communicate technical and non-technical trade-offs clearly in writing to an audience that includes both engineers and non-engineers. **Preferred** - Production experience with GPU-backed model serving, including autoscaling behavior under concurrent load, cold-start management and cost control. - Experience with OpenTelemetry and an enterprise observability platform such as Dynatrace, including instrumenting AI workloads rather than only conventional services. - Experience with model serving efficiency techniques: quantization, parameter-efficient fine-tuning, distillation or inference optimization. - Experience with data governance and access control on a lakehouse platform, such as Databricks Unity Catalog. - Experience operating AI systems in a regulated industry, particularly financial services, including auditability, retention and access-control requirements. - Experience with agentic or multi-step AI systems in production, including tracing and debugging multi-turn behavior. - Hands-on experience with a managed agent or tool gateway, such as AWS Bedrock AgentCore Gateway, and with running MCP servers in a governed environment. - Familiarity with emerging agent interoperability and agent identity standards, including Agent2Agent (A2A) style agent-to-agent delegation, agent discovery and capability advertisement, and workload identity for non-human actors. **Knowledge and Skills** - Ability to communicate complex technical information, both verbal and written, to all levels, including senior leadership. - Ability to solve problems at the source by offering simple, working solutions. - Responds promptly and effectively to resolve incidents, tasks, and projects. - Demonstrated ability and motivation to teach others. - Ability to gain the trust of others and build solid relationships across and vertically throughout the organization. - Effectively prioritize and execute tasks in a high-pressure environment. **Target Compensation:** A competitive base salary range from $154,120 to $226,042. This position is eligible for an annual variable bonus of cash and equity, between 10-20%. Bonus amounts are based on individual performance. Final compensation within the range is influenced by many factors including role-specific skills, depth and experience level, industry background, relevant education and certifications. Candidates who reside in the following major metropolitan areas may be eligible for a premium on top of the posted range based on their specific zone: San Francisco, Seattle, Boston, New York City, Los Angeles and San Diego. INDENGLP #zip #LI-Remote **Benefits** - Excellent benefits package that includes 401(K) match, adoption assistance, parental leave, tuition reimbursement, comprehensive medical/ dental/vision and many nonstandard benefits that make us a Great Place to Work **Our Company Values:** To be successful in this role, Team Members need to be: - Positive by maintaining resiliency and focusing on solutions - Respectful by collaborating and actively listening - Insightful by cultivating innovation, accumulating business and role specific knowledge, demonstrating self-awareness and making quality decisions - Direct by effectively communicating and conveying courage - Earnest by taking accountability, applying feedback and effectively planning and priority setting **Expectations:** - Remain compliant with our policies processes and legal guidelines - All other duties as assigned - Attendance as required by department **Advice**! We understand that your career search may look different than others. Our hiring team wants to make sure that this would be a fit not just for us, but for you long term. If you are actively looking or starting to explore new opportunities, send us your application! **P.S**. We have great details around our stats, success, history and more. We're proud of our culture and are happy to share why - let's talk! Required degrees must have been earned at institutions of Higher Education which are accredited by the Council for Higher Education Accreditation or equivalent.

About Credit Acceptance Corporation

Credit Acceptance Corporation is a publicly traded company that provides automobile loans through indirect lending. The company operates in the United States and Canada. Credit Acceptance Corporation was founded in 1972 by Don Foss, and is headquartered in Southfield, Michigan. The company has been publicly traded since 1992.
Learn more about Credit Acceptance Corporation
Size
2,033 employees
Market Cap
$5.7 billion
Industry
Net Income
$421 million
Founded
1972
5 Year Trend
+13.9%
Revenue
$1.6 billion
NASDAQ

Similar Jobs

More Jobs at Credit Acceptance Corporation

More Information Technology Jobs

Find similar Staff Machine Learning Engineer, Platform (MLOPS) jobs: