Quick Summary
Transform messy literature into practical improvements. The research literature is vast, ambiguous, and constantly evolving. You will use your skills as a scientist to source, vet, implement,
Models are what they eat. But a large portion of training compute is wasted training on data that are already learned, irrelevant, or even harmful, leading to worse models that cost more to train and deploy.
At DatologyAI, we’ve built a state of the art data curation suite to automatically curate and optimize petabytes of data to create the best possible training data for your models. Training on curated data can dramatically reduce training time and cost (7-40x faster training depending on the use case), dramatically increase model performance as if you had trained on >10x more raw data without increasing the cost of training, and allow smaller models with fewer than half the parameters to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more details, check out our recent research on synthetic data scaling (BeyondWeb) and pretraining with domain-specific data (The Finetuner’s Fallacy).
We raised a total of $57.5M in two rounds, a Seed and Series A. Our investors include Felicis Ventures, Radical Ventures, Amplify Partners, Microsoft, Amazon, and AI visionaries like Geoff Hinton, Yann LeCun, Jeff Dean, and many others who deeply understand the importance and difficulty of identifying and optimizing the best possible training data for models. Our team has pioneered this frontier research area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make data curation easy for anyone who wants to train their own model on their own data.
This role is based in San Mateo, CA. We are in office 4 days a week.
The internship is for Summer 2027 and will take place sometime between May and August 2027.
About the Role
~1 min readAs a Research Intern at DatologyAI, you will conduct research investigating how intervention on training data can improve the quality and shape the behavior of deep learning models. Here is what your day-to-day would look like:
Transform messy literature into practical improvements. The research literature is vast, ambiguous, and constantly evolving. You will use your skills as a scientist to source, vet, implement, and improve promising ideas from the literature and your own creation.
Perform High-Risk, High-Reward Research. We want our interns to focus on problems that have massive potential to transform how data is ingested into future ML models. Rather than making incremental changes to current algorithms, we want you to work on novel project ideas that could change how we view data.
Conduct science driven by real-world needs. At DatologyAI, we understand that conference reviewers and academic benchmarks don’t always incentivize the most impactful research. Concrete customer needs and product improvements will guide your research.
Science is more than just experiments. We expect our Research Scientist Interns to collaborate closely with engineers, talk to customers, and shape the product vision
Ideal candidates should have strong coding skills with experience with one of the following:
We would like to hire students with practical experience and/or publications related to any of the following research topics:
Data research
Data pruning/curation
Curriculum learning
Synthetic data generation
Dataset distillation
Effects of training data on model behavior
Embedding models
Semantic search
Efficient ML
We would love to have you if you have practical experience and/or publications related to training large vision (especially video), language, and multimodal models.
Or teach us something new that you are passionate about that could improve data curation!
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- September 19, 2026
- First seen
- September 19, 2026
- Last seen
- September 19, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 52%
- Scored at
- September 19, 2026
Signal breakdown
Please let datologyai know you found this job on Jobera.
3 other jobs at datologyai
View all →Explore open roles at datologyai.
Similar Research Intern jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.