Staff PlatformOps Engineer (Emerging AI)

DAT

$198K — $246K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in software engineering, preferably as a senior or staff engineer managing production infrastructure.
  • 4+ years of experience building and maintaining AI platforms, successfully transitioning models from development to production.
  • Proficiency in languages like TypeScript/Node.js, Java, or Go, with a focus on creating user-friendly libraries.
  • Extensive AWS experience with production services including Kubernetes, Lambda, and infrastructure-as-code tools.
  • Hands-on experience in serving models at scale with techniques like GPU scheduling and batch processing.
  • Practical knowledge of LLM systems, including integration of vector databases and managed APIs like Amazon Bedrock.
  • Strong ability to implement comprehensive evaluation and monitoring for machine learning models.

Responsibilities

  • Design, build, and manage shared services for model deployment and operations.
  • Lead architectural direction for large-scale AI systems and ensure alignment with stakeholders.
  • Create and operate scalable AI infrastructure in AWS, incorporating various advanced tools.
  • Develop systems for evaluating model performance and improvements, utilizing advanced testing methodologies.
  • Manage inference costs and performance metrics, optimizing resource allocation effectively.
  • Implement safety and governance protocols for AI systems, ensuring compliance with security standards.
  • Drive best software development practices and mentor teams on production ML and distributed systems.

Benefits

  • Comprehensive medical, dental, and vision insurance.
  • Flexible vacation time, allowing for personal and family needs.
  • 401k matching with immediate vesting and employee stock purchase options.
  • Generous additional paid holidays, enhancing work-life balance.
  • Supportive programs for employee wellness, recognition, and assistance.
  • Free public transport pass for the Beaverton office, simplifying commute logistics.
  • Opportunity to contribute significantly to transformative projects in the trucking industry.
Full Job Description
Job Application Deadline: 10/31/2026

The Opportunity

As a Staff AI Platform Engineer, you'll build and own the platform that every AI and machine learning workload at DAT runs on. Freight is an uncertain business, and the models we ship reduce that uncertainty: rate forecasts, load-to-truck matching, document extraction, fraud signals, and the agentic workflows our brokers and carriers use to move freight faster. None of that reaches a customer without a platform that makes training, serving, evaluating, and monitoring models routine instead of heroic.

You'll set the architectural direction for that platform, from the model gateway and inference layer through feature and vector storage, evaluation harnesses, and production observability. This is a highly visible role for an engineer who wants their work multiplied across every AI team in the company.

What You'll Do
  • Platform Ownership: Design, build, and operate the shared services that engineers and business users use to ship models: a model gateway for LLM access, inference endpoints for real-time and batch scoring, feature storage, vector search, and a common SDK.
  • Technical Leadership: Lead architecture for large-scale AI systems, write the design documents, and drive alignment across Product, and Engineering and Business users on how models get built and shipped at DAT.
  • Cloud Architecture: Architect and run scalable, reliable AI infrastructure on AWS, including Bedrock, SageMaker, EKS, Redpanda, MSK (Kafka), Lambda, and S3, all defined in Pulumi..
  • Evaluation and Quality: Build the offline and online evaluation systems that tell us whether a model or prompt change is an improvement, including regression suites, LLM-as-judge pipelines, A/B and shadow testing, and drift detection.
  • Cost and Performance: Own inference cost and latency as first-class metrics. Right-size GPU and serverless capacity, tune batching, caching, and quantization, and give teams clear visibility into what their workloads cost.
  • Safety and Governance: Implement guardrails, prompt and output logging, PII handling, access controls, and model and dataset lineage so AI systems meet our security and customer data commitments.
  • Best Practices: Drive the adoption of modern software development practices across AI work, including automated testing, code reviews, CI/CD pipelines, and infrastructure-as-code.
  • Mentorship: Mentor engineers and data scientists on production ML and distributed systems, and raise the operational bar of every team that builds on the platform.
  • Incident Management: Lead the response and resolution for complex production incidents involving AI services, perform root cause analysis, and implement preventative measures.

The Skills and Experience You'll Bring
  • 10+ years of experience in software engineering, including significant time as a senior or staff engineer owning production infrastructure that other engineering teams depend on.
  • 4+ years building and operating machine learning or AI platforms, with models you've taken from notebook to production traffic and then kept healthy.
  • Strong development language skills ideally using TypeScript/Node.js, Java, or Go, with the software engineering discipline to build libraries other teams adopt willingly.
  • Extensive AWS experience with production systems, ideally including EKS/Kubernetes, Lambda, Redpanda, MSK/Kafka, S3, and Secrets Manager, plus infrastructure-as-code with Pulumi, Terraform or similar.
  • Hands-on experience serving models at scale: containerized inference, GPU scheduling, autoscaling, batching, and the tradeoffs between real-time, streaming, and batch scoring.
  • Practical experience with LLM systems in production, including retrieval-augmented generation, vector databases, prompt and context management, tool-calling or agent frameworks, and managed model APIs such as Amazon Bedrock.
  • Proven experience building evaluation and monitoring for models, not just services: offline eval sets, online metrics, drift and data quality checks, and a clear definition of what "regression" means for a model.
  • Strong background in event-driven and distributed systems using Redpanda, AWS MSK or Kafka
  • Fluency with observability and on-call health, designing dashboards, APM traces, logs, and alerts, defining SLOs, and using these tools to drive down incident frequency and MTTR.
  • Experience establishing and raising engineering standards for a team: PR and testing guidelines, deployment validation checklists, model release criteria, and post-incident review practices.
  • Demonstrated technical leadership: leading cross-team designs in ambiguous problem spaces, writing clear system design documents and operations runbooks, and driving alignment across Engineering, Product, and Operations.
  • Proven mentoring track record, especially helping engineers and data scientists grow in system design, observability, and operational excellence.
  • Excellent communication skills, with the ability to explain trade-offs and system behavior clearly to engineers, product managers, and non-technical stakeholders.

We'd be Extra Excited if You Have
  • Experience with model fine-tuning, distillation, or parameter-efficient training, and a clear point of view on when it beats prompting.
  • Background in marketplace, pricing, forecasting, or document extraction problems.
  • Freight, logistics, or supply chain domain experience.


  • Medical, Dental, Vision, Life, and AD&D insurance
  • Parental Leave
  • Flexible Vacation Time (FVT)
  • An additional 10 holidays of paid time off per calendar year
  • 401k matching (immediately vested)
  • Employee Stock Purchase Plan
  • Short- and Long-term disability sick leave
  • Flexible Spending Accounts
  • Health Savings Accounts
  • Employee Assistance Program
  • Additional programs - Employee Referral, Internal Recognition, and Wellness
  • Free TriMet transit pass (Beaverton Office)
  • Competitive salary and benefits package
  • Work on impactful projects in a cutting-edge environment
  • Collaborative and supportive team culture
  • Opportunity to make a real difference in the trucking industry
  • Employee Resource Groups


For Washington-based candidates, in compliance with the Washington State Pay Transparency Law, the salary range for this role is $198,000.00 - $246,000.00 + target bonus. DAT considers factors such as scope and responsibilities of the position, candidate's work experience, education and training, core skills, internal equity, and market and business elements when extending an offer.

#LI-RF1

#LI-hybrid

Similar Jobs

More Jobs at DAT

More Enterprise Technology Jobs

Find similar Staff PlatformOps Engineer (Emerging AI) jobs: