2+ years in AI safety, red teaming, trust & safety, adversarial ML, or security research from reputable organizations.
Hands-on adversarial experience such as jailbreaking models and penetration testing.
Understanding of safety work structures like harm taxonomies and evaluation designs.
Ability to navigate ambiguous situations and define problems before addressing them.
Strong writing skills for precise risk specification.
Passion for AI and interest in entrepreneurship.
Proven competitive success record.
Leadership and communication skills.
Proficiency in Python with capability to write production-quality code.
Responsibilities
Partner with safety and trust teams to address complex data issues and create revenue-driving programs.
Recruit and manage expert teams including security researchers and domain risk experts.
Design specifications for adversarial datasets and safety benchmarks used in frontier model evaluation.
Develop repeatable methodologies for human red team campaigns and automated attack generation.
Oversee projects from initial conversation to final delivery in a dynamic environment.
Support initiatives across multiple areas to enhance the safety practice as it scales.
Benefits
Opportunity to work at the forefront of AI safety and security.
Engagement with leading teams in trust and safety across AI labs.
Chance to influence the structure of safety methodologies in AI models.
Access to cutting-edge technology and tools in AI research.
Collaboration with top experts in the field, enhancing professional growth.
Full Job Description
Responsibilities
Project and client management: Partner directly with the safety, alignment, and trust & safety teams at frontier AI labs - scoping their hardest data problems, translating them into deliverable programs, and driving revenue.
Manage expert teams: Recruit and lead specialized contributors - red teamers, security researchers, trust & safety practitioners, and domain risk experts - and hold a high bar on the quality of what they produce.
Shape how frontier models are made safe: Design the specifications behind adversarial datasets, jailbreak and attack taxonomies, refusal-boundary and over-refusal sets, and safety benchmarks that labs use to measure and improve model behavior.
Build the red teaming machinery: Stand up repeatable methodologies - human red team campaigns, automated attack generation, coverage tracking against harm taxonomies - rather than one-off deliverables.
End-to-end execution: Own projects from first conversation through final delivery, in a fast-moving environment where the spec often does not exist yet.
Operational impact: Support initiatives across building, analysis, coordination, and execution as we scale the safety practice.
Required Qualifications
2+ years in AI safety, red teaming, trust & safety, adversarial ML, or security research; at a frontier AI lab, FAANG, top security firm, or equivalent.
Hands-on adversarial experience: jailbreaking or stress-testing frontier models, prompt injection research, offensive security, penetration testing, bug bounty, CTFs, or trust & safety investigations and enforcement.
Fluency with how safety work is actually structured (harm taxonomies, evaluation design, safety frameworks and model policy).
High agency and the ability to execute in ambiguity; comfort defining the problem before solving it.
Strong writing. Much of this job is specifying risk precisely enough that fifty other people can execute against it.
Genuine passion for AI and interest in entrepreneurship.
Demonstrated track record of competitive success.
Strong leadership and communication skills.
Python proficiency and the ability to write production-quality code.
Preferred Qualifications
Published safety, alignment, or security research.
Track record in bug bounty, CTF, or model red teaming competitions.
Experience building automated red teaming or eval pipelines.
Depth in a high-consequence risk domain (cyber, bio, fraud and financial crime, child safety, influence operations).
Experience using AI tools to automate workflows and build products.