About this roleSpeak cares deeply about learners actually learning and improving with Speak. We have a dedicated Proficiency team to own how we measure learning efficacy and speaking proficiency, from unit-level mastery checks to standalone proficiency tests to onboarding placement, all in service of understanding users' proficiency levels and learning gains in an accurate, transparent and actionable manner. Because everything happens remotely and asynchronously in the app, keeping scores fair, stable over time, and resistant to gaming is both hard and genuinely interesting.
We're looking for an Assessment Design Lead: the person who defines what we measure, why, how, and signs off on whether our assessments actually measure the right thing. You will be staffed on the Proficiency team and report to the Head of Learning Design and Curriculum. If solving reliable, at-scale speaking assessment excites you and you want to directly shape Speak's efficacy story, we'd love to hear from you.
What you'll be doing- Own Assessment Design - Define what Speak measures, why, and how often across three distinct assessment types (Curriculum Mastery Assessment, Proficiency Test, Placement Test) - with the Proficiency Test as the immediate focus, expanding to the other two as the pod's priorities evolve. Keep constructs (fluency, pronunciation, grammar, task achievement) clearly separated and each aligned to CEFR or a comparable speaking proficiency standard, so no single assessment conflates domains it wasn't designed to measure.
- Define Constructs & Build the Rubric/Blueprint Layer - Translate fuzzy goals like "measure fluency" or "measure pronunciation" into concrete, scoreable constructs, item blueprints, and rubrics that an item writer can generate items against and an ML Engineer can build a grading model against.
- Own Validity & the Quality Bar - Sign off on content validity for every assessment that ships. Decide what "mastery" or a passing score operationally means, catch cases where an assessment is measuring the wrong thing before it ships, own the rubric/rater guidelines behind the human-labeled data our ML scoring models are evaluated against, and audit items/rubrics for bias across learner subgroups. Making sure scores stay comparable as the assessment evolves and stay meaningful against attempts to game an unproctored test.
- Design and run the validity evidence plan - so validity is built into the process rather than checked only after launch. This includes concurrent/criterion studies benchmarking Speak's assessments against external proficiency measures (CEFR-anchored exams, expert human ratings), so we can say what a Speak Score means in terms the outside world already trusts.
- Partner Tightly with Product and ML - Work closely with the Product Manager and ML Engineers on automated scoring, calibration, and feedback generation. You own the construct and quality bar, they own the model. Neither works without the other, and the loop between you is the product.
What we're looking forMust-Haves- Assessment/Psychometric Design: 4+ years designing rubrics, blueprints, and item specs for a real, shipped language assessment product (or equivalent depth in closely related psychometric/measurement work) - not just academic theory. Can explain reliability and validity in plain language and knows how to catch a test that's measuring the wrong construct.
- Language Proficiency Domain Expertise: Deep familiarity with frameworks like CEFR (or ACTFL, IELTS/TOEFL band descriptors) and what separates "did you learn what we taught you" from "how good is your speaking overall."
- Fairness Across Learner Populations: Can identify whether an item or rubric unfairly penalizes specific L1 backgrounds or accents (differential item functioning) - essential for a speech-based test serving learners across dozens of native languages.
- Translates Qualitative Technical: Can turn a construct like "pronunciation quality" into something concrete enough for an ML engineer to build a scoring pipeline against, without either oversimplifying or getting lost in academic nuance.
- Quantitative Rigor: Comfortable running or interpreting the statistics behind a rubric or rater system - inter-rater reliability (e.g., Cohen's/Fleiss' kappa), classical test theory, and basic IRT concepts - enough to know whether a scoring system is actually reliable, not just plausible.
- Ownership of Quality Bar: Comfortable being the sign-off authority on content validity - makes the call clearly and follows through on it, rather than deferring to data alone or product pressure to ship.
- AI Fluency & Judgment: Uses AI tools directly in their own workflow (e.g., drafting item variants, testing rubric language, exploring construct definitions) and has real judgment about when AI-generated output is precise enough to ship vs. needs a human rewrite - distinct from spec'ing work for the ML Engineer to build.
- Comfort with Ambiguity: Comfortable operating in a 0-to-1 environment. Can wear multiple hats, take a fuzzy goal and turn it into a concrete plan, communicate tradeoffs clearly, and keep momentum without waiting for perfect clarity or team setup
Nice-to-Haves- Speech/pronunciation science background - can own pronunciation frameworks and L2-specific error taxonomy directly
- Familiarity with adaptive testing or IRT-adjacent concepts (even if not the primary psychometrician)
- Experience at a large-scale language testing organization or similar high-rigor assessment environment
- Experience thriving in an EdTech startup environment, especially in a newly forming team or 0-to-1 mandate
- Advanced degree (Master's or PhD) in psychometrics, measurement, applied linguistics, SLA, or a related quantitative field; track record of shipped assessment work still matters more
- Has authored technical/validity reports or published assessment research
How We WorkThis role is designed to be highly collaborative with our Assessment ML Engineer. Success depends on a tight loop where constructs, rubrics, and model outputs co-evolve together - from the earliest fuzzy construct through to a shipped scoring model - rather than a one-time handoff from a finished spec to an implementation.