Linguist II

PRI Global

$80K — $95K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Linguistics, Computational Linguistics, Computer Science, Speech Science, or related field
  • 1+ years of experience in Linguistics, Language Technologies, NLP, or AI/ML data operations
  • Native or near-native fluency in English and at least one additional language
  • Knowledge of syntax, semantics, pragmatics, sociolinguistics, and corpus linguistics
  • Familiarity with Large Language Models (LLMs) and their applications
  • Experience with speech and text data in multiple languages
  • Comfortable working in a fast-paced, collaborative environment

Responsibilities

  • Apply linguistic expertise to support LLM and generative AI systems
  • Collaborate on data collection, curation, annotation, and localization efforts
  • Contribute to annotation schemas and guidelines for LLM training data
  • Evaluate and quality-check datasets for language models
  • Support methods for generating synthetic annotated data
  • Assist in model evaluation and linguistic error analysis
  • Contribute to AI safety through linguistic review of model outputs

Benefits

  • Opportunities for professional growth and skill development
  • Collaborative work environment with cross-functional teams
  • Engagement in cutting-edge AI and language technology projects
  • Focus on ethical and responsible AI practices
  • Contribution to innovative product development in AI space
Full Job Description
Description: Linguist II - Summary
We are looking for a linguist to help develop language components for AI-powered products, including large language models (LLMs) and voice-enabled technologies. We are seeking candidates with solid linguistic data analysis skills, programming familiarity, and language technology experience to contribute to data collection, synthetic data generation, and annotation tasks in support of LLM/AI training, evaluation, alignment, and AI agent development.

Job Responsibilities
Apply linguistic expertise in syntax, semantics, pragmatics, and sociolinguistics to support LLM and generative AI systems
Collaborate with linguists, data operations teams, and ML engineers on data collection, curation, annotation, and localization efforts for model training and fine-tuning
Contribute to the development and maintenance of annotation schemas and guidelines for LLM training data (e.g., instruction-tuning, preference labeling, RLHF)
Evaluate and quality-check datasets used for pre-training, fine-tuning, and alignment of language models
Support the development of programmatic methods for generating synthetic annotated data at scale
Assist in model evaluation efforts including prompt-based testing, red-teaming, and linguistic error analysis
Contribute to AI safety and responsible AI practices through linguistic review of model outputs (e.g., hallucination detection, bias identification, tone and pragmatic appropriateness)
Participate in experiments to assess data quality, annotation consistency, and downstream model performance

Basic Qualifications
Bachelor's degree in Linguistics, Computational Linguistics, Computer Science, Speech Science, or related field
1+ years of experience in Linguistics, Language Technologies, NLP, or AI/ML data operations (or equivalent)
Native or near-native fluency in English and at least one additional language
Knowledge of syntax, semantics, pragmatics, sociolinguistics, corpus linguistics, and other areas of linguistics
Familiarity with Large Language Models (LLMs), their applications and data practices (training data, evaluation, prompting, fine-tuning)
Exposure to LLM evaluation methodologies (human evaluation, automated metrics, adversarial testing)
Experience working with semantic ontologies, taxonomies, or intent/slot frameworks
Proficiency using AI Agents/Chatbots
Experience with database queries and data analysis processes (SQL, spreadsheets, R, Unix, or others)
Experience working with speech and text data in multiple languages
Comfortable working in a fast-paced, highly collaborative environment with evolving priorities

Preferred Qualifications
Master's degree in Linguistics, Computational Linguistics, Language Technologies, or a related field
Familiarity with machine learning frameworks, NLP libraries, and tools (e.g., Hugging Face, spaCy, NLTK, PyTorch)
Exposure to statistical language modeling or training data pipelines
Strong organizational skills and attention to detail

Custom Fields:
Name: Compelling Story & Candidate Value Proposition
Value: See RSD

Name: Does this role deal with graphic and/or objectionable content to any extent?
Value: No

Name: How will performance be measured
Value: See RSD

Name: Good to have skills
Value: See RSD

Name: Top 3 must-have HARD skills
Value: See RSD

Name: Typical Day in the Role
Value: See RSD

Name: Team Name
Value: None

Name: If yes, approximately how often might CW's engaged by your team encounter graphic or objectionable content (imagery, video, or written) in the course of their work with your team?
Value: None

Name: Story Behind the Need - Business Group & Key Projects
Value: See RSD

Similar Jobs

More Jobs at PRI Global

More Information Technology Jobs

Find similar Linguist II jobs: