Job Description:The Postdoctoral Associate will address safety and transparency gaps in our understanding of AI companions (e.g., risks like self-harm and deceptive/unwarranted anthropomorphization).
More specifically, the postdoc will research the extraction and application of character trait vectors for inference time steering. This involves techniques in the area of mechanistic interpretation. E.g., monitoring neuron activation patterns across model layers when the system is prompted with datasets designed to induce specific traits like philosophical archetypes. By identifying these activation signatures, we can develop steering mechanisms to influence model behavior at inference time. While inference time steering has a per-query cost vs. the one time cost of fine tuning, in the case of AI companions it allows for essentially infinite combinations of specific personality "facets" and how much individual facets should be scaled on a per-end-user basis, without the cost of any additional training.
Objectives:- How character traits are represented across different architectures. For example, dense vs. Mixture of Experts architectures, where specific experts might diverge in trait expression, resulting in unexpected (and potentially dangerous) behavior with certain types of interactions.
- Understanding how complexity and stability of character traits evolve with respect to model parameter count. For example, some has shown that larger models can exhibit emerging mechanisms to "resist" alignment or even hide certain behavior when in a testing environment.
- Developing methods for efficient computation of character traits. There is a per-query cost for inference time steering, regardless of its advantages over fine tuning for AI companions. As part of this project, the fellow will explore mechanisms to reduce these costs.
The Postdoctoral Associate will:- Design and implement experiments focused on understanding and modifying the behavior of Generative AI models.
- Lead projects related to AI Companions and mentor graduate and undergraduate students.
- Prepare manuscripts for submission to top Computer Science venues.
Requirements: - PhD in Computer Science or closely related field.
- Strong research record demonstrated by publications in top AI/NLP/Security venues.
- Research experience working with Generative AI applications and internals.
- Programming experience (Rust, C++, Python, and Cuda preferred).
- Ability to conduct independent research and work collaboratively with graduate students.
Additional Information:Salary: $85,000 annually. Note: Visa sponsorship is not available for this position.
Binghamton University is a tobacco-free campus. Application Instructions:All applicants must apply through this Hirezon/Interview Exchange system and should submit:
- Resume/CV
- Cover letter which addresses the specific reasons for your interest in this opportunity.
- Names and contact information of two individuals who can provide recommendations
You may add additional files/documents after uploading your resume/vitae.
After filling out the contact information, you will be directed to the upload page.
Returning Applicants - Login to your The Research Foundation for SUNY at Binghamton Careers Account to check your completed application.
Please contact us if you need assistance applying through this website.
URL: www.binghamton.edu/computer-science/