Socure Inc.

Staff Data Scientist, Watchlist

Socure Inc. • $135K — $160K *
Finance & Insurance
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Master's or PhD in Computer Science, Computational Linguistics, Statistics, Applied Mathematics, or related field; or equivalent experience.
  • 7+ years of experience in data science or machine learning, specifically with NLP, entity resolution, or information extraction.
  • Strong preference for experience in AML, sanctions screening, or financial crime detection.
  • Hands-on experience building NLP pipelines for entity extraction and named entity recognition at scale.
  • Strong proficiency in Python and major ML libraries, including PyTorch, spaCy, and HuggingFace Transformers.
  • Proficient in SQL and experience with large-scale data pipelines.
  • Excellent communicator able to articulate complex technical concepts to non-technical stakeholders.

Responsibilities

  • Enhance the quality and timeliness of data through advanced ingestion pipelines.
  • Conduct data quality analysis to identify issues and ensure high-quality inputs for models.
  • Utilize AI and NLP to enrich raw source data into structured formats.
  • Develop NLP systems for entity consolidation and resolving identities.
  • Measure and improve entity resolution accuracy and coverage.
  • Design advanced NLP models for real-time matching across diverse unstructured data sources.
  • Create models that help customers optimize their screening thresholds.

Benefits

  • Opportunity to lead technical initiatives and collaborate with product and engineering teams.
  • Direct impact on a critical product for AML compliance in global finance.
  • Mentorship opportunities to foster a culture of technical excellence.
  • Engagement with cutting-edge NLP advancements and real-world applications.
Full Job Description
ABOUT THE ROLE

We are looking for a Staff Data Scientist to join Socure's Watchlist Data Science team. Watchlist sits at the heart of global AML compliance - our platform screens hundreds of millions of entities in real time across sanctions lists, PEP databases, and adverse media sources for banks, fintechs, and payment companies worldwide.

As a Staff Data Scientist, you will work on the hardest problems in entity matching and classification: scaling our patented real-time matching engine, building advanced Natural Language Processing (NLP) models for Named Entity Recognition (NER) and Information Extraction, and bringing next-generation research to production. This is a senior individual contributor role with broad technical ownership and direct impact on a product that helps the world's financial institutions manage sanctions and AML risk.

WHAT YOU'LL DO

Data Quality & Enrichment

  • Improve the quality, coverage, and freshness of Watchlist's underlying data through next-generation ingestion pipelines.
  • Design and execute rigorous data quality analysis pipelines to identify anomalies, evaluate dataset health, and ensure high-fidelity inputs for downstream model training.
  • Apply NLP and AI to classify and enrich raw source data into normalized schemas - extracting structured entity attributes from unstructured sanctions, PEP, adverse media, and enforcement sources.
  • Expand multilingual capabilities to support global screening across Latin and non-Latin scripts.

Entity Resolution

  • Build and improve NLP systems that consolidate how watchlist identities are represented. Developing Information Extraction and Named Entity Recognition (NER) pipeline to deduplicate entities across lists and resolve aliases into canonical profiles..
  • Develop approaches to handle how entity profiles change over time as names, aliases, and sanctions status evolve.
  • Measure and benchmark entity resolution quality, driving continuous improvement in coverage and accuracy.


Match Engine & Risk Scoring

  • Design and scale advanced NLP models and algorithms that perform real-time name matching and identity classification across diverse, multilingual unstructured data sources.
  • Build multi-signal risk scoring that combines name similarity, entity type, geography, list type, and other attributes into unified, calibrated risk scores.
  • Maintain and improve benchmarking frameworks, golden datasets, and regression tests that keep the match engine at the highest levels of recall and precision.


Analytics, Tuning & Evaluation

  • Build models and analytics that help customers tune their screening thresholds to the right operating point for their risk appetite and entity mix.
  • Develop backtesting and counterfactual analysis capabilities so customers and internal teams can understand how model or threshold changes would affect screening outcomes.
  • Design evaluation frameworks for AI-powered autonomous decision systems - defining correct behavior, calibrating confidence thresholds, and monitoring for drift in production.


AML Risk Detection

  • As Watchlist expands into payment screening, build the mathematical analysis and feature engineering needed to detect AML risk patterns across transaction data and payment message fields.
  • Develop and maintain the AML taxonomy and risk signal library that underlies Watchlist's classification and detection capabilities.
  • Apply graph-based methods to surface indirect risk exposure - identifying entities connected to sanctions risk even when they are not directly listed.


Research & Technical Leadership

  • Lead technical initiatives across Watchlist Data Science and shape the team's long-term approach to entity matching, enrichment, and AI.
  • Collaborate closely with Product and Engineering to translate research into production-grade systems at scale.
  • Stay current with advances in NLP, large language models, and entity resolution; prototype and deploy relevant techniques (e.g., advanced NER, LLM-based extraction) to AML use cases.
  • Mentor peers and contribute to a culture of technical rigor and continuous improvement.


WHAT YOU BRING

  • Master's or PhD in Computer Science, Computational Linguistics, Statistics, Applied Mathematics, or a related field; or equivalent professional experience.
  • 7+ years of experience in data science or machine learning, with meaningful work in NLP, entity resolution, or information extraction.
  • Experience in AML, sanctions screening, adverse media, or financial crime detection is strongly preferred.
  • Hands-on experience building and deploying NLP pipelines for entity extraction, named entity recognition, and record linkage at production scale.
  • Familiarity with multilingual NLP and non-Latin script processing is a strong plus.
  • Experience with LLMs and agentic AI frameworks (e.g., LangChain/LangGraph) is a plus.
  • Strong proficiency in Python and major ML libraries (PyTorch, spaCy, HuggingFace Transformers).
  • Strong SQL proficiency and experience with large-scale data pipelines and production ML systems.
  • Excellent communication skills - able to translate model performance tradeoffs into compliance and business language for non-technical audiences.


Note: We cannot provide Sponsorship at this time.

You must be located in one of our talent hubs: New York, San Francisco, Seattle, or Miami.

Follow Us!

YouTube | LinkedIn | X (Twitter) | Facebook

About Socure Inc.

Socure is a New York-based technology company that provides digital identity verification services. The company's products use artificial intelligence and machine learning to verify the identities of individuals in real-time. Socure's customers include financial institutions, online marketplaces, and other businesses that need to verify the identities of their users. The company was founded in 2012 by Sunil Madhu and Johnny Ayers and has raised over $70 million in funding to date.
Learn more about Socure Inc.
Size
250 employees
Industry
Founded
2012

Similar Jobs

More Jobs at Socure Inc.

More Finance & Insurance Jobs

Find similar Staff Data Scientist, Watchlist jobs: