Archer Data Scientist

Archer Technologies LLC

$129K — $215K *
US-AnywhereRemote in Livermore, CA
Legal & Accounting
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in data science, ML engineering, or AI-driven software development.
  • Strong programming skills in Python (NumPy, Pandas, PyTorch/TensorFlow, LangChain, or equivalent).
  • Experience with vector databases and retrieval systems (Pinecone, FAISS, Weaviate, etc.).
  • Hands-on experience with RAG pipelines, embedding models, and LLM orchestration (OpenAI, Bedrock, Hugging Face, etc.).
  • Solid understanding of data pipelines, ETL frameworks, and cloud-native deployment on AWS.
  • Familiarity with Elasticsearch, PostgreSQL, and API integration patterns.
  • Knowledge of ML lifecycle management, including model training, evaluation, and monitoring.

Responsibilities

  • Design, train, and evaluate LLM-based pipelines for document understanding and regulatory reasoning.
  • Implement and optimize RAG architectures for semantic retrieval.
  • Develop and maintain model fine-tuning workflows and knowledge distillation.
  • Collaborate with ML Ops teams to integrate AI models into production-ready APIs on AWS.
  • Ensure data lineage and version control across AI and analytics pipelines.
  • Build and manage end-to-end data pipelines for legal and compliance data.
  • Architect and maintain intelligent Knowledge Bases to support AI-driven search and compliance reasoning.

Benefits

  • Collaborative work environment with cross-disciplinary teams.
  • Opportunity to work with cutting-edge AI technologies in LegalTech and RegTech.
  • Support for continuous learning and professional development initiatives.
  • Flexible working arrangements, including remote options.
  • Access to advanced tools and infrastructure in AWS cloud.
Full Job Description
Data Scientist - LLM & Data Pipeline Engineering (LegalTech / RegTech AI)

Overview:

We are seeking an experienced Data Scientist with a strong background in AI model integration, data pipeline development, and knowledge base (KB) engineering to support our next-generation LegalTech / RegTech AI platform.

This role blends applied machine learning, data engineering, and software development, focusing on building scalable pipelines that connect large language models (LLMs) to structured and unstructured data through retrieval-augmented generation (RAG) and vector database architectures.

The ideal candidate is passionate about operationalizing AI - from training and fine-tuning models to deploying intelligent retrieval systems in AWS cloud environments.

Key Responsibilities

1. AI Model Integration & Development

  • Design, train, and evaluate LLM-based pipelines for document understanding, obligation extraction, and regulatory reasoning.
  • Implement and optimize RAG architectures, combining LLMs with vector databases for semantic retrieval.
  • Develop and maintain model fine-tuning workflows, embedding generation, and knowledge distillation.
  • Collaborate with ML Ops teams to integrate AI models into production-ready APIs and services on AWS.
  • Measure and improve model precision, recall, latency, and interpretability.


1.5 Agentic and MCP Knowledge Integration:

  • Design and maintain agentic multi-component processes (MCPs) that enable context-aware reasoning across multiple data sources and agents.
  • Implement AI agents capable of dynamic tool use, autonomous task decomposition, and multi-context knowledge retrieval.
  • Develop pipelines that support agent memory, self-reflection, and knowledge synthesis across distributed systems and knowledge bases.
  • Collaborate with engineering teams to integrate MCP-driven agents with retrieval, analytics, and workflow orchestration layers, ensuring compliance with regulatory reasoning frameworks.


2. Data Pipeline Engineering

  • Build and manage end-to-end data pipelines for ingestion, transformation, embedding, and indexing of legal and compliance data.
  • Orchestrate data workflows leveraging AWS services (e.g., S3, Lambda, Glue, SageMaker, Step Functions, RDS).
  • Develop scalable ETL/ELT processes to feed both relational (PostgreSQL) and vector databases (e.g., Pinecone, FAISS, Weaviate, Elastic Vector Search).
  • Ensure data lineage, reproducibility, and version control across AI and analytics pipelines.
  • Automate retraining and evaluation pipelines for continuous learning from user feedback.


3. Knowledge Base & Information Retrieval

  • Architect and maintain intelligent Knowledge Bases (KBs) to support AI-driven search, summarization, and compliance reasoning.
  • Implement advanced retrieval techniques using ElasticSearch / Elastic Vector Search and embedding-based retrieval.
  • Align KB structures with business ontologies and regulatory taxonomies to support explainable AI outputs.
  • Collaborate with domain experts and PMs to enrich KB metadata and enhance model context relevance.


4. AWS & Deployment

  • Deploy and scale AI pipelines using AWS services such as SageMaker, Lambda, ECS/EKS, API Gateway, and CloudFormation/Terraform.
  • Implement model and data monitoring solutions for drift detection, latency management, and cost optimization.
  • Collaborate with DevOps to maintain secure, reliable, and compliant cloud environments.


5. Cross-Functional Collaboration

  • Partner with engineering, product, and compliance teams to align AI models with regulatory and data governance requirements.
  • Work closely with QA and Professional Services teams to validate AI outputs and improve client-facing performance.
  • Document architectures, experiment results, and data flows to ensure transparency and reproducibility.


Preferred Experience

  • Experience building AI products for LegalTech, RegTech, or compliance automation.
  • Familiarity with agentic AI frameworks (e.g., OpenAI MCP, CrewAI, LangGraph, or AutoGen).
  • Background in document intelligence systems, multi-agent orchestration, or knowledge graph integration.
  • Experience with LangChain, LlamaIndex, or similar frameworks for RAG orchestration.
  • Hands-on knowledge of MLOps tools and data versioning (DVC, MLflow, Weights & Biases).
  • Understanding of governance, interpretability, and ethical AI


Qualifications

  • 5+ years of experience in data science, ML engineering, or AI-driven software development.
  • Strong programming skills in Python (NumPy, Pandas, PyTorch/TensorFlow, LangChain, or equivalent).
  • Experience with vector databases and retrieval systems (Pinecone, FAISS, Weaviate, Qdrant, or Elastic Vector Search).
  • Hands-on experience with RAG pipelines, embedding models, and LLM orchestration (OpenAI, Bedrock, Hugging Face, etc.).
  • Solid understanding of data pipelines, ETL frameworks, and cloud-native deployment on AWS.
  • Familiarity with Elasticsearch, PostgreSQL, and API integration patterns.
  • Knowledge of ML lifecycle management, including model training, evaluation, and monitoring.


Soft Skills

  • Strong problem-solving and system design capabilities.
  • Excellent communication skills for cross-disciplinary collaboration.
  • Passion for structured documentation, reproducibility, and experimentation.
  • Adaptable mindset with focus on performance, scalability, and reliability.


Success Indicators

  • Scalable and well-documented RAG pipelines supporting production of AI workloads.
  • High model accuracy, retrievability, and latency efficiency.
  • Reliable data flow from ingestion to inference with minimal manual intervention.
  • Increased explainability and compliance assurance across AI outputs.


Additional Information:

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice at management discretion based on business need.

Pay Transparency Notice: We're committed to fair and transparent pay practices. In line with state pay transparency laws, for positions in this location, we offer a base pay of $129,200.00 to $215,400.00 USD, plus benefits. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Please contact our Talent Acquisition team at [redacted] for the range and related compensation details. Actual pay may vary based on location, experience, skills, and internal equity.

Similar Jobs

More Jobs at Archer Technologies LLC

  • Demand Generation Manager
    $88K — $146K *
    Remote
    Business Services
    Remote in United States
  • Full Stack Engineer
    $100K — $140K *
    Livermore, CA 94550 (Alameda County)
    Information Technology
    In-Person
  • Archer Data Scientist
    $129K — $215K *
    Remote
    Legal & Accounting
    Remote in Livermore, CA

More Legal & Accounting Jobs

Find similar Archer Data Scientist jobs: