specter
specter11mo ago
New

Software Engineer - ML Infrastructure

United StatesUnited States·San Franciscofull-timemid
Software EngineerSoftware Engineering
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Designing and implementing scalable ML training and inference pipelines for perception models (object detection, tracking, classification, segmentation) and VLMs.

Requirements Summary

Strong software engineering fundamentals in Python, with working proficiency in a systems language (Rust, Go, C++) for performance-sensitive data and inference paths.

Technical Tools
Software EngineerSoftware Engineering

Responsibilities

~1 min read
  • →

    Designing and implementing scalable ML training and inference pipelines for perception models (object detection, tracking, classification, segmentation) and VLMs.

  • →

    Developing continuous training and evaluation systems to improve model performance from production data feedback loops.

  • →

    Designing large-scale multi-modal data pipelines for ingesting, processing, and indexing video and sensor data spanning both batch and streaming workloads.

  • →

    Creating data pipelines for ingesting, labeling, versioning, and managing massive multi-modal sensor datasets (video, radar, lidar, thermal).

  • →

    Implementing model monitoring, A/B testing frameworks, and performance analytics for deployed perception systems.

  • →

    Collaborating with perception researchers to transition models from research to production at scale across thousands of edge nodes.

  • →

    Building tools and infrastructure for distributed training, hyperparameter optimization, and experiment tracking.

Requirements

~1 min read
  • Strong software engineering fundamentals in Python, with working proficiency in a systems language (Rust, Go, C++) for performance-sensitive data and inference paths.

  • Working proficiency with ML frameworks (PyTorch, TensorFlow) and model optimization tooling.

  • Deep experience building and operating model inference systems at scale, request routing, batching, autoscaling, caching, and latency/throughput tuning under real production load.

  • Hands-on experience with distributed compute frameworks for ML and data workloads (Ray, Spark, or equivalent), including GPU cluster management and orchestration.

  • Strong understanding of distributed systems fundamentals: partitioning, replication, backpressure, exactly-once semantics.

  • Experience with vector databases (QDrant, LanceDB, or equivalent) for similarity search and retrieval workloads.

  • Familiarity with LLM/VLM serving frameworks (VLLM, SGLang, TensorRT-LLM) in production is a strong plus.

  • Familiarity with video processing, sensor fusion, or multi-modal perception systems is a plus.

Location & Eligibility

Where is the job
San Francisco, United States
On-site at the office
Who can apply
US

Listing Details

Posted
October 3, 2025
First seen
September 25, 2026
Last seen
September 26, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
19%
Scored at
September 26, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

specterSoftware Engineer - ML Infrastructure