New
USD 120000-135000/yr

Data/ML Engineer

mid
Machine Learning EngineerData
2 views0 saves0 applied

Quick Summary

Key Responsibilities

Work closely with the Accelerator leadership team to align data engineering and machine learning initiatives with overarching goals and long-term vision.

Requirements Summary

Work effectively in a modern, professional software and data engineering environment with a strong understanding of Agile concepts and practices.

Technical Tools
Machine Learning EngineerData

The Accelerator seeks a Data/ML Engineer to strengthen our data team and advance the engineering, enrichment, and provisioning of the data we collect.

 

The Accelerator at Princeton includes a portfolio of multiple planned independent and intersecting tools, built on a shared data and compute platform serving computational social scientists at research institutions across North America, Europe, and Africa. The Data/ML Engineer will work within our team to help drive data engineering and machine learning initiatives and collaborations. They will play a crucial role in building and operating the pipelines that transform large-scale social media and web behavior data into research-ready data products, and in developing the machine learning and enrichment capabilities that extend their value. They will work on problems that have no precedent and little source material, requiring novel solutions. They will also be responsible for working with the other teams within the Accelerator and our external partners to help foster collaboration and create an impactful environment for our users.

 

Responsibilities

~1 min read

What We Offer

~1 min read
✓Design, build, and operate data pipelines across the Accelerator's medallion architecture, with end-to-end ownership of transformation layers that serve researchers.
✓Ensure the accuracy, integrity, and quality of data to be made available through the Accelerator, including data quality validation at pipeline boundaries and enforcement of versioned schema contracts.
✓Diagnose and optimize distributed data processing workloads at production scale.
✓Develop deployment automation, CI/CD, and release processes for data products, including versioned data releases and researcher-facing change documentation.
  • Design, develop, and operate ML and NLP enrichment pipelines over large-scale text and behavioral data, including language identification, translation, and topic and content classification.
  • Own the full lifecycle of enrichment models: selection, evaluation against labeled data, batch inference architecture, cost efficiency, and reprocessing and versioning strategy.
  • Develop ML-ready feature layers and data products to support advanced research use cases.
  • Evaluate and apply large language model workflows and other emerging AI methods where they demonstrably improve outcomes, with attention to their validity for downstream scientific analysis.
  • Apply statistical analysis and modeling to characterize datasets, estimate coverage, and support research design.

 

  • Contribute to cost attribution, visibility, and governance across institutional workspaces, including cluster policies, budget controls, and storage lifecycle management.
  • Design data and ML workloads to operate within the platform's cost governance framework.
  • Develop automation for workspace and project provisioning as institutions and research projects onboard.
  • Operate within Unity Catalog governance, multi-tenant isolation, and research data security requirements.

 

  • Work effectively in a modern, professional software and data engineering environment with a strong understanding of Agile concepts and practices.
  • Modern Software Engineering Foundations: agile (Scrum), DevOps, CI/CD, code review, and pair programming, with working knowledge of cloud compute platforms to support collaborative, scalable, and efficient development.
  • Author and maintain researcher-facing documentation and provide direct technical support to research users of the platform.
  • Collaborate with research teams to define data products, sampling frames, and enrichment requirements, and apply state-of-the-art techniques to ongoing scientific challenges.
  • Stay current with the latest advancements in data engineering, machine learning, and relevant fields to continuously innovate.
  • Build strong relationships with external partners, driving collaborations that enhance the Accelerator's scientific impact.

Requirements

~2 min read
  • 3+ years of relevant experience as a data engineer, machine learning engineer, or data scientist, which may include graduate research and internship experience, with a record of building production systems that operate reliably at scale. Experience working in a remote, agile environment.
  • Bachelor's degree or equivalent in a relevant field.
  • Strong proficiency in Python and SQL, and hands-on experience with distributed data processing (e.g., Apache Spark) on large data volumes.
  • Experience building, evaluating, and operating machine learning or NLP pipelines, including batch inference.
  • Working knowledge of cloud data platforms.
  • Strong communication and interpersonal skills to effectively collaborate with researchers in the field, other engineers at various levels of experience, and administrative and leadership team members.

 

  • Experience with Azure and Databricks, including Unity Catalog.
  • Experience with infrastructure-as-code (e.g., Terraform), containers, and CI/CD tooling.
  • Experience with large-scale social media, web behavior, or text-as-data research.
  • Familiarity with large language model annotation workflows and their evaluation.
  • Publications in reputable scientific journals or conferences is desirable.

 

Princeton University is an Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability status, protected veteran status, or any other characteristic protected by law.

 

The University considers factors such as (but not limited to) scope and responsibilities of the position, candidate's qualifications, work experience, education/training, key skills, market, collective bargaining agreements as applicable, and organizational considerations when extending an offer. The posted salary range represents the University's good faith and reasonable estimate for a full-time position; salaries for part-time positions are pro-rated accordingly.

 

If the salary range on the posted position shows an hourly rate, this is the baseline; the actual hourly rate may be higher, depending on the position and factors listed above.

 

The University also offers a comprehensive benefit program to eligible employees. Please see this link for more information.

36.25
No
180 days
No
No
No
Mid-Senior Level
#Ll-DP1
$120,000 to $135,000

Location & Eligibility

Where is the job
—
Location terms not specified

Listing Details

Posted
October 7, 2026
First seen
October 7, 2026
Last seen
October 7, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
55%
Scored at
October 7, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Data/ML EngineerUSD 120000-135000