Location: This is a hybrid remote/in-office role.
Our team:This position is for the
Data Nexus team within Study Experience Engineering, focused on building Medidata's next generation of
AI-enabled data platforms and intelligent applications. The team works across
Generative AI, Agentic workflows, AWS, Snowflake, Apache Iceberg, Airflow, Python, and Java to power scalable data ingestion, transformation, APIs, analytics, and automation. You will work in a highly collaborative environment with engineering, product, QA, PMO, and platform teams to solve complex data and application challenges and bring AI-driven capabilities into Medidata's clinical technology ecosystem.
Who we are looking for:Reporting to the Director, Engineering, the
Senior AI & Data Engineer - Study Experience will design, build, and operate production-grade AI and data solutions that enable intelligent applications, automation, analytics, and operational workflows.
This is a hands-on engineering role with a primary focus on Generative AI and agentic applications, supported by strong data and application engineering capabilities. You will build AI-enabled services, connect enterprise data and applications to LLM-powered workflows, and develop scalable data pipelines using modern cloud technologies.
What You'll Do- Design and build production Generative AI applications, AI agents, and LLM-powered workflows for enterprise use cases.
- Develop AI services and backend capabilities using Python, APIs, and cloud-native engineering patterns.
- Build RAG and retrieval workflows that securely connect enterprise data and applications to AI solutions.
- Implement agentic workflows involving tool calling, structured outputs, retrieval, workflow execution, and external system integrations.
- Work with AI orchestration frameworks and emerging standards such as LangGraph, LangChain, MCP, or similar technologies.
- Implement testing, evaluation, guardrails, monitoring, and observability to improve the quality and reliability of production AI solutions.
- Design and maintain scalable data pipelines and ETL/ELT workflows that support AI, analytics, and operational applications.
- Build data solutions using AWS, Snowflake, Python, SQL, and modern data engineering technologies.
- Develop and integrate APIs and backend services, including Java-based applications where needed.
- Apply strong software engineering practices including automated testing, CI/CD, Git, security, documentation, and production support.
- Use AI-assisted development tools and coding agents as part of day-to-day design, coding, testing, debugging, and code review.
- Collaborate with engineers, architects, product teams, and business stakeholders on technical designs and production solutions.
- Mentor engineers through technical guidance, code reviews, and knowledge sharing.
What We're Looking For- 5+ years of professional experience in software engineering, AI engineering, data engineering, or a related technical discipline.
- Strong hands-on experience developing production applications and services using Python.
- Practical experience building or integrating Generative AI applications, LLM workflows, agentic systems, or AI-enabled services.
- Good understanding of RAG, embeddings, vector search, retrieval, grounding, and LLM orchestration concepts.
- Experience with at least one AI orchestration framework or emerging protocol such as LangGraph, LangChain, MCP, or equivalent technologies.
- Strong SQL and data engineering fundamentals, including experience building production data pipelines and integrations.
- Experience working with AWS and modern data platforms such as Snowflake.
- Experience developing or integrating REST APIs and backend services.
- Strong understanding of modern software engineering practices including automated testing, CI/CD, version control, observability, and production support.
- Ability to independently troubleshoot complex issues across AI services, APIs, data pipelines, and cloud environments.
- Strong communication and collaboration skills and the ability to contribute effectively to technical design discussions.
Preferred Skills- Experience with managed AI platforms or models such as AWS Bedrock, Anthropic Claude, OpenAI APIs, Azure OpenAI, or similar technologies.
- Experience with vector and search technologies such as OpenSearch, pgVector, Pinecone, Weaviate, or similar platforms.
- Experience building MCP servers, tools, or enterprise AI integrations.
- Familiarity with data technologies such as Airflow, Kafka, Apache Iceberg, or dbt.
- Experience with Docker, Kubernetes, Terraform, or similar cloud-native technologies.
- Familiarity with AI evaluation, prompt/version management, structured logging, and AI observability.
- Understanding of AI governance, data privacy, security, responsible AI, and human-in-the-loop patterns.
- Experience working in regulated industries such as Life Sciences or Healthcare.
As with all roles, Medidata sets ranges based on a number of factors including function, level, candidate expertise and experience, and geographic location.
- The salary range for positions that will be physically based in the NYC Metro Area is $114,750-153,000
- The salary range for positions that will be physically based in all other locations within the United States is $102,750-137,000.
Base pay is one part of the Total Rewards that Medidata provides to compensate and recognize employees for their work. Most sales positions are eligible for a commission on the terms of applicable plan documents, and many of Medidata's non-sales positions are eligible for annual bonuses. Medidata believes that benefits should connect you to the support you need when it matters most and provides benefits, including medical, dental, life and disability insurance, 401(k) matching, family leave, flexible paid time off; and 10 paid holidays per year.
#LI-MM1
#LI-Hybrid