We'd love for you to join Basis.
Rapid advances in frontier AI are shaping how billions of people work, learn, & live every day. Much of this advancement is due to scaling (more compute + more/better data better coverage over desired distributions). But, while compute spend is growing up to 4x per year, publicly available data increases by only 1.03x per year! That gap becomes even starker when moving beyond LLMs into voice, video, & the physical world. Hence, there is desperate demand for high quality, high diversity, & carefully curated recordings of human interaction and expression.
Leading labs require vast amounts of human feedback, behavior, culture, speech, & judgement. We're already working with leading labs and have one of the largest contributor networks in the industry, to provide exactly that.
What you'll do (Software Engineer):
- Build pipelines that score, filter, and rank millions of user recordings each day using in-house ML-models and custom fraud detection systems
- Design scalable systems from scratch that automatically control millions of dollars in user rewards and ad spend.
- Work directly with leading labs, assess their needs, & deliver high-quality and heavily filtered recordings for their frontier models.
- Collaborate closely with founders in a fast-iteration environment, independently owning systems end-to-end with high autonomy.
You: Enjoys system design; is very intentional & sweats the details; moves very fast with AI coding tools; loves building; thrives on solving problems no one has ever solved before.
What you'll do (Research Engineer):
- Train state-of-the-art internal models, public benchmarks, and technical publications using the one-of-a-kind multi-modal data we collect.
- Collaborate with top researchers at leading labs to create extremely challenging open-source benchmarks to push the frontier of voice and interaction modeling.
- Example: frontier speech-to-text models still struggle with regional dialects like Moroccan Arabic. We collect tens of thousands of hours of hand-transcribed Moroccan Arabic audio and use it to train internal models that accelerate transcription and automatically verify data quality.
- Example: we collect tens of millions of preference ratings to train in-house reward models used in RL pipelines, for both enterprise clients and open-source releases/benchmarking.
You: deep conceptual understanding of deep learning; experience fine-tuning strong open models; ideally with publication track record.
What you'll do (MTS/Research-Software Hybrid Role):
Employment Details:
- Competitive compensation + equity
- 20 days PTO
- 33A Company retreat
- Platinum-class medical, dental & vision insurance
- Lunch & dinner provided every work day (weekdays, 10-8)
- Beautiful office space in San Francisco
- Best-in-class immigration support
- 49A Most importantly, the ability to work on some of the most consequential technical challenges of our time.