17h ago
New

Internship - Spatial AI and Navigation

United KingdomUnited Kingdom·Londonfull-timeentry
OtherInternship
2 views0 saves0 applied

Quick Summary

Overview

Here at Humanoid, we believe in a future where robots amplify human potential. That’s why we’ve set out on a mission to build the world’s most capable, commercially-scalable, and safe humanoid robots.

Technical Tools
OtherInternship

Here at Humanoid, we believe in a future where robots amplify human potential. That’s why we’ve set out on a mission to build the world’s most capable, commercially-scalable, and safe humanoid robots. We’re bringing that mission to life with HMND‑01 - our rapidly developed humanoid platform being deployed in real industrial environments - and we’re growing the team to take it even further.

We’re building software systems that enable robots to operate effectively in the real world expanding human capability and redefining how work gets done.

 

We’re looking for interns who are curious, proactive, and excited to work on real-world robotic systems.

This is an open-ended internship, you won’t be confined to a single component, but will work across perception, navigation, and multimodal systems, collaborating closely with the team to find where you can have the most impact.

You may work anywhere along the stack, from camera systems (timestamping, synchronization, validation), through perception and scene understanding, to navigation and integration with locomotion. The scope is intentionally broad. We’re looking for people who are excited to dive into unfamiliar areas and learn quickly.

This is a full-time internship (5 days per week) over the winter, based in our London Euston office, where you’ll contribute to real systems from early on with guidance and support from experienced researchers and engineers.

Duration: 12 weeks | Start date: Nov | Compensation: Competitive pay + we'll keep you fed (seriously, the food is good)

 
  • Develop perception systems for robot navigation and interaction in real-world environments

  • Work on focused problems within Vision-Language(-Action) or multimodal models (components, datasets, evaluation)

  • Run and analyse experiments using existing pipelines

  • Improve data quality through curation and labeling

  • Explore scene understanding, 3D perception, or navigation methods and apply them to real systems

  • Prototype ideas and iterate quickly with guidance

  • Collaborate on integrating models into robotic platforms

     
  • Pursuing a degree in Computer Science, Machine Learning, Robotics, or a related field.

  • Strong foundations in machine learning and/or computer vision.

  • Hands-on experience with PyTorch and training ML models.

  • Experience running experiments and interpreting results.

  • Interest in multimodal models, 3D vision, spatial reasoning, navigation or embodied AI.

  • Ability to take ownership and iterate with guidance.

  • Strong problem-solving skills and attention to detail.

  • Fast learner, comfortable in a research-driven, fast-moving environment.

 

Complete the challenge below and submit your solution as a public GitHub repository — include a README with instructions to run your system, example outputs, and a short note on your design choices. You will be able to include your GitHub repository URL when you fill out the application form, alongside your name and CV.

We’re not looking for standard solutions, we're looking for how you think. The strongest submissions are creative, original, and push beyond the obvious.

Build a system that compares two phone recordings of the same small indoor area, such as a room, and identifies meaningful changes in 3D: moved furniture, new obstacles, or removed objects.

Robots need to understand how their surroundings change over time. Your system should identify what changed and where it happened, while distinguishing actual changes from differences in camera viewpoint, occlusion, or incomplete observations.

Capture two short videos using your phone, moving, adding, or removing a few objects between recordings. Record from different viewpoints, with enough overlap to compare the same space.

At a minimum, your system should:

  • Generate a shared 3D representation of the scene from the two recordings

  • Identify and visualize objects or regions that were moved, added, or removed in 3D

  • Describe the detected changes using semantic labels or natural language

There are no constraints on real-time performance, We’re intentionally leaving the approach open, use any tools, models, frameworks, or agentic workflows you find effective.

  • A working codebase

  • Clear instructions on how to run your system

  • Example input(s) and output(s)

  • (Optional) A short note explaining your design choices and tradeoffs

 
  • Simplicity and usability of your solution

  • Creativity in approach

  • Quality of spatial change detection

  • Clear, compelling presentation of results

  • Coherence between detected changes, geometry, and semantics

Make something you’re proud of!

Location & Eligibility

Where is the job
London, United Kingdom
On-site at the office
Who can apply
GB

Listing Details

Posted
October 8, 2026
First seen
October 8, 2026
Last seen
October 8, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
65%
Scored at
October 8, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Employees
125
Founded
2009
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Internship - Spatial AI and Navigation