Offensive AI Engineer - Frontier AI Security & TestingWho we are looking forState Street Global Cybersecurity (GCS) is looking for a hands-on Offensive AI Engineer to join Frontier AI Security & Testing (FAST). FAST is a new research and development team within Regulatory Assurance, Penetration Testing & Offensive Research. This role works at the intersection of frontier AI, cloud engineering, data engineering and offensive security. You will design, deploy and run the secure AWS-based platforms, evaluation harnesses and data pipelines that let State Street use frontier AI models safely as an offensive security capability.
This is an engineering role, not an advisory or architecture-only one. You will build, configure, deploy, troubleshoot and automate. You will work directly with frontier models, agentic frameworks and enterprise APIs. The work also includes putting technical controls in place so that AI is used in an auditable, governed way in a highly regulated financial services environment.
What you will be responsible forAs an Offensive AI Engineer you will:
AWS AI infrastructure and deployment
- Design, engineer and deploy secure, isolated AWS environments to host and access frontier AI models. This includes Amazon Bedrock, SageMaker, private endpoint connectivity (PrivateLink/VPC endpoints), EKS/ECS, Lambda and API Gateway.
- Build and maintain infrastructure-as-code (Terraform or CloudFormation) and CI/CD pipelines for AI workloads, with repeatable and auditable deployment patterns.
- Engineer network segmentation, identity and access controls, secrets management and egress restrictions that create enforceable trust boundaries around agentic AI activity.
Frontier model engineering and evaluation
- Integrate with frontier model provider APIs and SDKs from major commercial vendors. Build model-agnostic abstraction layers that allow fast model onboarding and side-by-side comparison.
- Design and build standardized evaluation harnesses and benchmarking frameworks. These should measure model capability, precision, false-positive rate, reliability, cost and failure modes across offensive security tasks in a repeatable way.
- Develop agentic workflows, tool-use integrations and multi-agent orchestration patterns (e.g., MCP, LangGraph, Strands, or similar) that support reconnaissance, attack-path analysis and adversary simulation in controlled lab environments.
- Stress-test models against known AI failure modes, including prompt injection, jailbreak susceptibility, data leakage, tool misuse, scope drift and unsafe emergent behavior.
Data science and Databricks
- Design and build Databricks pipelines (PySpark/SQL) to ingest, curate and analyze model telemetry, evaluation results and offensive security data sets.
- Apply data science methods to evaluation results, including statistical comparison, scoring methodology, trend analysis and regression detection across model versions.
- Build data products and dashboards that give leadership evidence of model performance, cost and risk.
API integration
- Build secure, scalable integrations between AI platforms and enterprise systems, security tools and data sources using REST APIs, event-driven architectures and tool-execution brokers.
- Develop reusable API patterns and connectors for tool invocation, context injection and output validation.
AI controls, logging and governance
- Implement technical controls that limit agent autonomy where needed. These include guardrails, human-in-the-loop checkpoints, rate and scope limits, kill-switch mechanisms and least-privilege tool access.
- Engineer full logging, monitoring and observability that capture agent inputs, outputs, tool calls and decision paths to support auditability and post-incident analysis.
- Align AI use with enterprise Responsible AI, model risk management, legal and compliance requirements. Produce technical evidence suitable for audit and regulatory review.
- Document architectures, control implementations and evaluation methods to a standard that holds up under internal and external scrutiny.
What we valueThese skills will help you succeed in this role:
- A builder's mindset: you are comfortable taking an idea from prototype to a hardened, production-grade capability.
- Deep hands-on experience deploying AI-enabled applications on AWS, particularly Amazon Bedrock, SageMaker, Lambda, API Gateway, EKS/ECS, IAM, KMS and VPC networking.
- Direct experience working with frontier LLMs through vendor APIs, including prompt and context engineering, tool use and function calling, agentic workflows and model evaluation.
- Strong Python skills, plus experience with GenAI and agentic frameworks (LangChain, LangGraph, Strands, CrewAI, Pydantic, MCP, or similar).
- Hands-on Databricks experience (PySpark, SQL, Delta Lake, workflows) and a working grounding in data science and statistical analysis.
- Experience designing and consuming REST APIs and event-driven integrations.
- A working understanding of AI security risks and controls, such as the OWASP Top 10 for LLM Applications, MITRE ATLAS and the NIST AI RMF.
- Familiarity with offensive security concepts, adversary tactics and frameworks such as MITRE ATT&CK (strongly preferred).
- Strong critical thinking and problem-solving skills, and the ability to explain complex technical results clearly to engineers, leadership and risk partners.
- The ability to research, learn and apply new models, tools and techniques quickly as the frontier AI landscape changes.
Education & Preferred Qualifications- Bachelor's degree in Computer Science, Engineering, Data Science, Information Systems, Artificial Intelligence, or equivalent practical experience.
- 5+ years of hands-on experience in software, cloud, data or security engineering.
- 2+ years of hands-on experience building and deploying GenAI, LLM or agentic AI solutions on cloud platforms, preferably AWS.
- Experience with infrastructure-as-code (Terraform preferred) and CI/CD tooling.
- AWS certifications (e.g., Solutions Architect, Machine Learning Specialty, Security Specialty) and/or Databricks certifications are a plus.
- Offensive security certifications (e.g., OSCP, OSEP, GPEN) or penetration testing experience are a plus.
- Experience in financial services or another highly regulated industry is preferred.
Salary Range: $120,000 - $202,500 Annual
The range quoted above applies to the role in the primary location specified. If the candidate would ultimately work outside of the primary location above, the applicable range could differ.
Employees are eligible to participate in State Street's comprehensive benefits program, which includes: our retirement savings plan (401K) with company match; insurance coverage including basic life, medical, dental, vision, long-term disability, and other optional additional coverages; paid-time off including vacation, sick leave, short term disability, and family care responsibilities; access to our Employee Assistance Program; incentive compensation including eligibility for annual performance-based awards (excluding certain sales roles subject to sales incentive plans); and, eligibility for certain tax advantaged savings plans.
For a full overview, visit https://hrportal.ehr.com/statestreet/Home.