6h ago
From $150/yr

Senior Data Scientist AI Evaluation

CanadaCanadaRemoteFull-timesenior
Data ScientistData
2 views0 saves0 applied

Quick Summary

Requirements Summary

Approximately 6–10 years of experience in quantitative data science, machine learning, or a related technical discipline, with focused experience in measurement, evaluation, experimentation,

Technical Tools
Data ScientistData

We are seeking a senior-level data scientist to establish and advance the evaluation practices that measure the quality, reliability, and correctness of AI models and intelligent agents. In this role, you will transform complex quality questions into measurable standards, trusted ground truth, and rigorous evaluation frameworks. You will design repeatable testing processes that identify performance regressions and support safer, more confident AI releases. Working closely with Product, Engineering, Analytics Engineering, and business stakeholders, you will turn evaluation results into practical improvements. This is a high-impact individual contributor position offering ownership of a developing AI evaluation practice within a fast-growing financial technology environment. You will help shape how AI quality is measured, validated, documented, and continuously improved across the organization.

 

Requirements

~2 min read
  • Approximately 6–10 years of experience in quantitative data science, machine learning, or a related technical discipline, with focused experience in measurement, evaluation, experimentation, or model validation.

  • Strong quantitative and statistical foundations, including experience designing rigorous experiments and interpreting results with appropriate uncertainty measures.

  • Demonstrated experience defining meaningful evaluation metrics, establishing ground truth, and assessing the quality of complex or ambiguous model outputs.

  • Proficiency in Python and SQL, with experience applying these skills to data analysis, model assessment, and evaluation workflows.

  • Experience evaluating machine learning models in production environments and translating findings into practical improvements.

  • Ability to validate automated graders against human judgments and account for variability, non-determinism, and potential measurement bias.

  • Strong problem-solving and analytical judgment, particularly when developing evaluation approaches for new or evolving AI systems.

  • Excellent communication skills, with the ability to explain evaluation results, trade-offs, and recommendations to technical teams and business stakeholders.

  • Proven ability to collaborate across functions while independently owning complex analytical projects in a fast-paced environment.

  • A quantitative degree in data science, statistics, mathematics, computer science, or a related field is an advantage; equivalent practical experience is also welcome.

  • Hands-on experience evaluating large language models (LLMs) or AI agents in production, including evaluation harnesses, LLM-as-judge calibration, and continuous integration regression gates, is a plus.

  • Experience evaluating text-to-SQL systems, analytics agents, or other AI applications where outputs can be verified against underlying data is an advantage.

  • Background in fintech, brokerage, financial services, or other domains where incorrect outputs can create significant business or risk consequences is desirable.

  • Familiarity with AI tools used in research, analysis, and engineering workflows is beneficial.

What We Offer

~2 min read
✓Competitive compensation: Salary package designed to reflect experience and expertise.
✓Stock options: Opportunity to participate in the company's long-term growth.
✓Health benefits: Benefits designed to support employee health and well-being.
✓Home-office setup allowance: One-time allowance of USD $500 to support your remote workspace.
✓Monthly stipend: USD $150 per month provided through a Brex card.
✓Remote work environment: Opportunity to work remotely within the Americas, with Canada as the target location for this listing.
✓Technical ownership: Lead the development of evaluation practices and establish standards that influence AI quality across the organization.
✓Cross-functional impact: Work closely with product, engineering, analytics, and business teams to turn rigorous measurement into better AI systems.
✓Professional development: Build expertise in AI evaluation, model validation, and the responsible deployment of intelligent systems in a growing technology environment.

Location & Eligibility

Where is the job
Canada
Remote within one country
Who can apply
CA

Listing Details

Posted
October 9, 2026
First seen
October 9, 2026
Last seen
October 9, 2026

Posting Health

Days active
-1
Repost count
0
Trust Level
80%
Scored at
October 9, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Senior Data Scientist AI EvaluationFrom $0k