Description
Job Description - Senior Forge AI Engineer
About The Position:
Most AI engineering roles ask you to bolt an LLM onto an existing product. This one doesn't. We're building Forge - a greenfield agentic platform that replaces the human-driven SDLC with an autonomous agent chain, end-to-end. You'd be working across the full stack: orchestration, multi-agent coordination, MCP server design, RAG pipelines, eval harnesses, and the safety primitives that make non-deterministic systems trustworthy in production. Not one layer. All of it.
The platform is live and shipping. The hard problems are real: agent memory and state across multi-step runs, inter-agent trust and failure recovery, cost attribution at scale, compliance in a regulated industry. If you want to work at the frontier of agentic systems - not on a demo, on something that actually runs - this is that role.
Key Responsibilities:
• Design and build agents across the Forge chain - orchestration, tool-use surfaces, MCP server integrations, prompt and RAG pipelines - and set the agentic design patterns the team ships against.
• Own multi-agent coordination architecture: inter-agent trust boundaries, message schema contracts, partial failure handling, and recovery across the agent graph.
• Build and evolve the MCP server layer - tool schema design, capability contracts, versioning - so Forge's tool-use surface is well-defined and safe to extend.
• Design agent memory and state management: context window strategy, long-term memory integration, state persistence across complex multi-step workflows.
• Own agent quality end-to-end: eval harnesses, behavioral benchmarks, continuous evaluation pipelines, and the acceptance criteria that gate every agent release.
• Instrument agent-specific observability - token tracing, tool-call chain logging, cost attribution per run - so the platform is debuggable and defensible.
• Build safe-by-design primitives: least-privilege tool-use, PII/PHI handling, structured output validation, red-team partnership with Security.
Required Qualifications
• 6+ years of software engineering experience, with at least 2 years building and operating LLM-powered or agentic systems in production - systems that ran, failed, and got fixed.
• Production experience with at least one agentic framework (LangGraph, CrewAI, Autogen, Semantic Kernel, MCP, or comparable) and a clear understanding of the trade-offs each introduces.
• Hands-on experience authoring MCP servers or designing tool-use APIs with explicit capability contracts and versioning.
• Experience with agent memory architectures - context management, long-term memory stores, state persistence across multi-step workflows.
• Strong fundamentals in Python and/or TypeScript; able to own production agent code from design through operation.
• Proficient in prompt engineering, tool/function calling, and multi-step orchestration - and fluent in the failure modes of non-deterministic systems.
• Working knowledge of RAG patterns, vector databases, and LLM evaluation methodologies applied to production agent outputs.
• Working knowledge of AI/ML safety engineering: OWASP LLM Top 10, prompt injection defenses, RAG hardening, adversarial agent testing.
Preferred Qualifications
• Pharma, life sciences, or healthcare data experience is a plus - not a requirement.