Yale University

Data Engineer, Temporary

Yale University$68K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Master's degree in a scientific discipline or equivalent experience.
  • 3 years relevant experience in research or data processing.
  • Proficiency in Python for data processing and building pipelines.
  • Ability to evaluate computational methods and define accuracy measures.
  • Experience working with error-prone, incomplete, or inconsistent data.

Responsibilities

  • Test and evaluate approaches for entity extraction from machine-generated text.
  • Design and run structured comparisons of extraction methods.
  • Analyze results and assess reliability of derived data for text type.
  • Document methods, results, and contribute to methodological practice.
  • Participate in project meetings and present findings to stakeholders.

Benefits

  • Access to a collaborative research environment at Yale Library.
  • Opportunity to engage in cutting-edge experimentation with AI and machine-generated text.
  • Experience in the cultural heritage sector focusing on digital collections.
  • Flexible working arrangements in a scholarly community.
Full Job Description
Overview

Position Focus: (3000 Characters):

This is a time-limited research position investigating how descriptive metadata might be derived from machine-generated text produced by handwritten text recognition (HTR) and optical character recognition (OCR) applied to digitized library and archival materials. The appointment has a fixed end date and is exploratory. It is not a production role. Work produced during the appointment is not expected to be loaded into any system of record.

The position is a partnership between the Metadata Services Unit, part of Resource Discovery Services, and Digital Collections and Access.

Yale Library is currently experimenting with AI models applied to machine-generated text produced as a byproduct of digitization. How that text can be turned into structured, reliable description is an open question. Current descriptive practice does not capture information at this level of granularity. It is not known which extraction approaches are viable, how accurate they are, what output they produce, or how that output could be represented in a catalog or archival management system. The purpose of this position is to investigate those questions and document what is found, including negative results.

Work will be conducted against exported machine-generated text carrying a known error rate rather than human-corrected text, so findings reflect realistic conditions. All source material is in the public domain.

Main responsibilities include:

Entity extraction and resolution: testing and evaluation (50%)

Identify and test approaches to extracting and resolving entities and concepts from machine-generated text, including names, roles, dates, and places. Design and run structured comparisons across approaches. Establish evaluation criteria, measure accuracy, and characterize error patterns. Build and run Python-based processing pipelines and, where appropriate, use models hosted on Yale computing infrastructure.

Analysis and processing of descriptive output (30%)

Characterize the results of extraction and resolution and assess their reliability by material and text type. Analyze how the resulting data would and would not map to current and emerging bibliographic and archival description standards and to systems such as Alma and ArchivesSpace. Compare extracted values against existing catalog and finding aid records. Evaluate and test semantic storage solutions for project data, including graph and vector databases. Identify where human review would be required before any output could be relied upon.

Documentation and reporting (20%)

Document methods, code, evaluation criteria, and results so the work is reproducible by others. Contribute to documentation of methodological practice. Prepare written summaries of findings and open questions and present results to project stakeholders. Participate in project meetings.

Work is carried out under the direction of the supervisor, in consultation with staff in Resource Discovery Services and Digital Collections and Access.

Required Skills and Abilities

1. Demonstrated proficiency in Python for data processing and analysis, including building and running pipelines to extract, clean, and transform data.

2. Demonstrated ability to design and carry out structured evaluations of computational methods, including defining accuracy measures, benchmarking results, and testing reproducibility.

3. Demonstrated ability to work with incomplete, inconsistent, or erroneous source data, identify patterns in errors, and assess the relative reliability of derived values.

4. Demonstrated ability to work independently with limited supervision on open-ended problems, exercise judgment where no established process or procedure exists, and adapt approach based on findings.

5. Demonstrated ability to document technical methods and results clearly and accurately for both technical and non-technical audiences.

Preferred Education, Experience and Skills: (400 Characters):

Experience with text recognition, OCR, HTR, or natural language processing tools. Experience working in high-performance computing environments. Familiarity with bibliographic or archival description standards. Interest in cultural heritage collections.

Physical Requirements: (1000 Characters):

Sustained concentration at a computer workstation for extended periods. Ability to sit for sustained periods.

Principal Responsibilities

1. Designs, plans, implements and evaluates complex multifaceted research projects/systems, including method development, adaptation and validation. 2. Collaborates with PI to define research endeavors and the development of research hypothesis and approach. 3. Solves complex methodology, protocol, procedural and research problems through design of techniques, procedures and policies that will achieve research goals. 4. Develops long range plans for supplies and equipment to ensure smooth operation of laboratory. 5. Investigates, analyzes and evaluates complex data, data collection systems and methods to reach scientific conclusions and ensures the integrity of research data. 6. Collaborates with PI to coordinate major lab renovations working with Facilities, Health and Safety, and Purchasing to ensure deadlines, safety standards and purchases are according to plans and within budgets. 7. Prepares scientific reports and papers for research proposals and published reports. 8. May perform other duties as assigned. Required Education and Experience Master27s Degree in a scientific discipline and three years of experience or an equivalent combination of education and experience.

Job Posting Date

09/02/2026

Job Category

Professional

Bargaining Unit

NON

Compensation Grade

Clinical & Research

Compensation Grade Profile

Research Associate 2 MS (24)

Salary Range

$68,000.00 - $120,500.00

Time Type

Full time

Duration Type

Temporary / Casual (Fixed Term)

Work Model

On-site

Note

Yale University is a tobacco-free campus.

About Yale University

Yale University is a private Ivy League research university in New Haven, Connecticut. Founded in 1701, it is the third-oldest institution of higher education in the United States. Yale has a diverse student body and offers undergraduate and graduate degree programs in a range of academic fields. The university is known for its strong liberal arts program, as well as its professional schools of law, business, and medicine. Yale has produced numerous notable alumni, including five U.S. Presidents and 20 Nobel laureates.
Learn more about Yale University
Size
13,433 employees
Industry
Founded
1701

Similar Jobs

More Jobs at Yale University

More Information Technology Jobs

Find similar Data Engineer, Temporary jobs: