Wiley
• $109K — $156K *Qualifications
Responsibilities
Benefits
Job Description:
About the Role:
We'rebuilding the systems that turn one of the world's largest scientific corpora intoresearchintelligence. That meansproductionNLP pipelines running over millions of journal articles, extracting entities, classifications, claim tuples, and summariesoptimizedfor use by downstream agentic applications.We'relooking for a senior data scientist to owndomain-specificcontentmodeling work end to end, from the eval set through the pipeline stage that ships it.
You'lljoin a small, senior team where data scientists own their models in production.You'llwrite the code, own the evaluations, ship the changes, and stay accountable for the outcomes. This is a hands-on role for someone who wants to see their models through to real usersin a rapidly evolving market.
Job Responsibilities:
Design and build NLP enrichment pipelines that extract entities, classifications, claims, and summaries from scientificfull-textat scale.
Compare NLP approaches to extraction and enrichment against LLM-basedapproaches, andpick the right tool for each task. That means putting traditional NLP (NER, sequence labeling, classification), embedding-based retrieval, LLM prompting, and fine-tuned smaller models on the same table, and defending each choice with evaluation, cost, and operational tradeoffs. This is a core part of the job, not an occasional exercise.
Own evaluation. Build the golden setsin consultation with SMEs and vendors, choose the metrics, and make productive tradeoffs between speed, quality, and cost.
Contribute to agentic AI application work: tool-using systems that reason over the enriched corpus, where your NLP and evaluation background will shape how the agent grounds and defends its answers.
Work directly with editors, product managers, and engineers. Bring the modeling perspective into productdecisions, andtranslate stakeholderpushbackinto concrete modeling work.
Required Qualifications:
Strong NLP background across modern (LLMs, transformers, embeddings, retrieval) and classical (NER, classification, sequence labeling) approaches.You'vebuilt evaluationsand learned from the results.
Cleanpython.You are comfortable in exploratory notebooks and production repositories, and an engineer taking over a modeling output from you has a good head start.
A habit of comparing approaches and choosing the right one for the task. You can defend "prompt a large LLM" and "train a small classifier on 2,000 labels" with equal seriousness, back the choice with an eval and a costestimate, andknow what to do when performance drifts.
Preferred Qualifications:
Experience working with scientific or scholarly text.
Familiarity with AWS (S3, Batch, Lambda, SageMaker) and Parquet or Iceberg data lake patterns.
Experience running LLMs underreal costand latency budgets in production.
Some exposure to agentic AI applications: tool use, multi-step reasoning, guardrails, and evaluation of trajectories rather than single-turn outputs.
Salary Range:
109,500.00 USD to 156,833.33 USD #LI-JG1Similar Jobs



More Jobs at Wiley
More Information Technology Jobs


