Data Scientist

Simpson Thacher and Bartlett LLP

• $150K — $175K *
Finance & Insurance
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in data science, mathematics, statistics, computer science, engineering, finance, or related field required.
  • Master's degree in data science, computer science, statistics, computational linguistics or engineering preferred.
  • Over 2 years of experience in data science, machine learning engineering, or artificial intelligence, or a PhD in a related field.
  • Expertise in statistical programming (Python, R) and databases (SQL, Pinecone).
  • Proven experience in developing and validating regression and classification models.

Responsibilities

  • Support legal teams in using data to drive decision-making and enhance client representations.
  • Collaborate with departments such as Finance and IT to develop data solutions that meet operational goals.
  • Develop regression and classification models using advanced data science methodologies.
  • Fine-tune and deploy pre-trained language models for various NLP tasks including summarization and question-answering.
  • Design document segmentation and embedding approaches for improved information retrieval.
  • Conduct quantitative research to identify patterns and anomalies within large data sets.
  • Create visually appealing reports that present quantitative insights effectively to non-technical stakeholders.

Benefits

  • Flexible hybrid working arrangement available.
  • Access to ongoing training and development in advanced data science methodologies.
  • Opportunity to work on cutting-edge AI projects within the legal realm.
  • Collaborative work environment with multidisciplinary teams.
Full Job Description
Job Summary & Objectives

Simpson Thacher's Data Scientist will play a key role in delivering insights to the Firm's leadership, legal practices, Knowledge Department and other administrative functions. To do so, this expert will use a mix of statistics, machine learning, deep learning and natural language processing methods to derive meaning, build models and create systems that capitalize on both structured and unstructured data. This position will play a crucial role in pushing the Firm's artificial intelligence (AI) efforts forward and will work on some of the most advanced projects in the legal industry.

The Data Scientist position sits within the Firm's Data Science Team, which is a part of the Firm's Knowledge & Innovation Department. This position will work closely with lawyers and legal support professionals across practices, as well as other technical resources within the Firm. This role requires creative problem solving, analytical rigor, technical skill and an appreciation for the business and practice of law.

Essential Job Duties & Responsibilities

  • Support legal teams and relevant operational staff in delivering on opportunities to use data to drive decision-making and improve the efficiency and effectiveness of the Firm's client representations
  • Collaborate with Firm functional departments (e.g., Finance, Talent, Business Development, IT) to analyze data and develop solutions to support operational objectives
  • Develop regression and classification models using established and emerging data science methodologies
  • Chain, fine-tune and deploy pre-trained language models (e.g., BERT, Llama, Qwen, etc.) to optimize performance on a range of NLP tasks, including text classification, named entity recognition, and generative tasks such as summarization, clause and document generation, and question-answer exchanges
  • Design and deploy document segmentation and embedding approaches to facilitate information retrieval and retrieval augmented generation (RAG)
  • Conduct advanced quantitative research, using machine learning (ML) and natural language processing (NLP) techniques to understand patterns in large volumes of data, identify relationships, detect data anomalies and classify data
  • Configure practice-specific AI workflows and language technologies, which may require complex pipelines, prompt engineering, prompt chaining, and text operations
  • Design and deploy highly visual reports and user interfaces that surface quantitative insights in forms that are fit-for-purpose, modern and easily accessible to non-technical business professionals
  • Stay current with the latest advancements in LLMs, NLP, deep learning and ML research, implementing cutting-edge techniques and incorporating them into production models as appropriate
  • Document development processes, codebase, and best practices to facilitate knowledge sharing and maintain a well-organized, reproducible environment
  • Partner with other technical resources to refine data pipelines for recurring classes of analysis and data-driven solutions
  • Handle projects on request under the direction of the CKIO and other executive staff


Education

  • A bachelor's degree required, preferably in data science, mathematics, statistics, computer science, engineering, finance or a related field
  • Master's degree in data science, computer science, statistics, computational linguistics or engineering preferred
  • Prior coursework in deep learning, natural language processing, or information retrieval a significant plus


Skills and Experience

  • 2+ years in a data science, machine learning engineering, artificial intelligence or equivalent role, or a PhD in a related field
  • Highly proficient with statistical programming (e.g., Python, R) and databases (e.g., SQL, Pinecone)
  • Proven experience developing and validating linear and non-linear regression and classification models
  • Expertise in data transformation, data science and visualization libraries (e.g., pandas, scikit-learn, matplotlib, Seaborn)
  • Experience with natural language processing and related libraries (e.g., Hugging Face's Transformers, spaCy, NLTK) preferred
  • Ability to design and develop object-oriented machine learning systems beyond Jupyter notebooks a plus
  • Solid understanding of deep learning frameworks such as TensorFlow or PyTorch a plus
  • Proficiency with version control systems such as Git or equivalent tools for code management and collaboration
  • Able to translate business problems to technical logic and practical solutions
  • Able to communicate complex results clearly to a non-technical audience
  • Proactively develops and maintains technical knowledge in emerging data science areas
  • Experience in the legal field is a significant plus


Salary Information

NY Only: The estimated base salary range for this position is $150,000 to $175,000 at the time of posting.

The actual salary offered will depend on a variety of factors, including without limitation, the qualifications of the individual applicant for the position, years of relevant experience, level of education attained, certifications or other professional licenses held, and if applicable, the location in which the applicant lives and/or from which they will be performing the job. This role is exempt meaning it is not overtime pay eligible.

Simpson Thacher will not sponsor applicants for work visas for this position.

#LI-Hybrid

Similar Jobs

  • Graphite.com
    Data Scientist
    $160K — $200K *
    Graphite.com
    New York, NY 10025 (New York County)
  • Member of Data Staff
    $150K — $300K *
    Simile AI, Inc
    New York, NY 10025 (New York County)
  • Signals Analyst - Mid to Experienced Level
    $105K — $192K *
    National Security Agency/Central Security Service
    Fort George G Meade, MD 20755 (Anne Arundel County)
  • CACI International
    Data Scientist--JIATF
    $86K — $181K *
    CACI International
    Alexandria, VA 22304 (Alexandria City County)
  • TeleTech
    Data Scientist
    $77K — $176K *
    TeleTech
    Washington, DC 20011 (District Of Columbia County)
  • TeleTech
    Data Scientist
    $99K — $225K *
    TeleTech
    Mclean, VA 22101 (Fairfax County)

More Jobs at Simpson Thacher and Bartlett LLP

More Finance & Insurance Jobs

Find similar Data Scientist jobs: