The RoleAs an Applied AI Data Analyst, you will bridge the gap between our massive, unstructured financial and operational datasets to our core R&D team. You will own the evaluation platform of our agentic workflows forecasting pipelines, as well as improve customer facing insights leading to mission-critical decisions.
This role is part of the Applied AI Research department, with options to advance to senior data science and AI/ML engineering roles.
Location: New York, NY (In-Person | Relocation supported)
What you'll do- Algorithm & LLM Evaluation: Analyze edge cases, model predictions, hallucination rates, and extraction accuracy for our LLM-driven document processing systems and agentic workflows.
- Root Cause Analysis: Dig deep into complex datasets (inconsistent financial documents, fragmented enterprise systems) to determine why an extraction or time-series model succeeded or failed.
- Golden Dataset Curation: Define, curate, and maintain gold-standard evaluation datasets ("golden sets") for automated backtesting of AI agents and demand forecasting models.
- Performance Monitoring & Evals: Build dashboards, data transformations, and evaluation harnesses to monitor system health, model drift, and AI performance metrics across release cycles.
- Cross-Functional Collaboration: Partner directly with applied AI engineers, product managers and business operations managers to turn qualitative real-world problem definitions into quantitative product defining insights.
What we're looking forRequired- Education: Graduate degree (M.S. or Ph.D.) in Computer Science, Data Science, Statistics, Mathematics, Physics, or a related STEM field.
- Experience: 2-4 years of experience in a quantitative data analyst or AI evaluation role, working closely alongside AI/ML R&D teams.
- Programming & Agentic AI: Proficiency in Python and hands-on works with AI/ML models
- Data Stack: Strong capabilities with SQL and enterprise data infrastructure (PostgreSQL, Snowflake, dbt, AWS).
- Analytical Mindset: Exceptional ability to translate noisy, unstructured document data into rigorous evaluation insights for engineers.
Nice to have- Experience with LLM observability, tracing, and eval frameworks (LangSmith, Arize AI, or TruLens).
- Proficiency with agentic workflows or multi-agent orchestration.
- Familiarity with demand forecasting, time-series models, or statistical modeling.
- Comfort in fast-moving, high-growth NYC startup environments.
Perks + Benefits- Equity - own a meaningful piece of the company you're helping build
- Fully paid health coverage through Aetna - we cover 100% of employee premiums
- Top-tier dental and vision coverage through Guardian
- Carrot Fertility Pro - comprehensive fertility and family-forming support
- 12 weeks of paid parental leave
- Unlimited PTO - plus regular 4-day holiday weekends we actually take
- 401(k) through Vestwell
- Paid relocation support - we'll help you make the move to NYC
- Fully equipped workspace from day one - laptop, monitor, keyboard, and a $200 stipend to personalize your setup
- Team perks - catered Friday lunches, team dinners, and unlimited coffee + snacks featuring products from the brands we work with