Senior Data Scientist AI Evaluation
Quick Summary
Define ground truth, metrics, and scoring methods for models and agents. Build repeatable eval loops: Track quality over time and catch regressions before release.
We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!
We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.
Responsibilities
~1 min read- →Design AI evaluations: Define ground truth, metrics, and scoring methods for models and agents.
- →Build repeatable eval loops: Track quality over time and catch regressions before release.
- →Partner on infrastructure: Work with engineering and analytics engineering to operationalize eval harnesses.
- →Drive iteration: Translate eval results into actionable recommendations for system improvements.
- →Establish quality standards: Set evaluation guidelines, documentation, and review practices.
- →Mentor and align: Foster evaluation best practices and build a culture of measurable AI quality across the team.
- Track record of quantitative measurement rigor (e.g., LLM/model evaluation, metric validation, or experimentation).
- Strong statistical and ML foundation—you treat evaluations as experiments (sample sizing, confidence intervals, significance, handling non-determinism) and validate automated graders against human ground truth.
- Proficiency in Python and SQL, with experience evaluating models in production environments.
- Strong judgment in defining quality metrics and ground truth for ambiguous outputs.
- Excellent communication and cross-functional collaboration skills to align technical teams and leadership.
- Strong problem-solving ability in fast-paced, greenfield environments.
- 6–10 years in quantitative data science or ML, with focused experience in measurement or evaluation. A quantitative degree is a plus; equivalent industry experience is equally welcome.
Nice to Have
~1 min read- Hands-on LLM/agent evaluation in production, including eval harnesses, LLM-as-judge calibration, and CI regression gates.
- Experience evaluating text-to-SQL, analytics agents, or other systems where correctness is verifiable against data.
- Background in fintech, brokerage, or other domains where a wrong answer has real business or risk consequences.
- Fluency with AI tools in research and engineering workflows.
- Competitive Salary & Stock Options
- Health Benefits
- New Hire Home-Office Setup: One-time USD $500
- Monthly Stipend: USD $150 per month via a Brex Card
Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.
Location & Eligibility
Listing Details
- Posted
- October 7, 2026
- First seen
- October 7, 2026
- Last seen
- October 7, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 75%
- Scored at
- October 7, 2026
Signal breakdown
Alpaca builds financial services APIs for everyone globally.
View company profileSimilar Data Scientist jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.
