Senior Backend Engineer: Machine Learning Infrastructure
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps.
As a Senior Backend Engineer, you’ll help build the infrastructure that enables machine learning teams to develop, deploy, and operate models reliably at scale.
You’ll own critical backend services, data pipelines, and platform capabilities supporting workloads ranging from traditional machine learning to LLMs.
The role combines distributed systems engineering, cloud infrastructure, API development, observability, and platform design.
You’ll work closely with ML and product teams to understand their needs and turn them into reusable, self-service infrastructure.
You’ll have significant autonomy to make technical decisions, evaluate trade-offs, and take solutions from design through continuous improvement.
The environment is remote-first, collaborative, and strongly oriented toward ownership, thoughtful engineering, and modern AI-assisted development.
This is an opportunity to have a direct impact on how teams build and operate machine learning products at scale.
- Design, build, deploy, and operate high-load distributed backend services and APIs that support machine learning infrastructure.
- Take end-to-end ownership of core ML services and associated data pipelines, from system design and implementation through deployment, observability, maintenance, and continuous improvement.
- Build reliable, scalable, and reusable infrastructure components that make machine learning workloads easier for product and ML teams to run and operate.
- Partner closely with ML and product engineers to understand their requirements and translate them into effective platform capabilities and services.
- Make and communicate technical decisions by evaluating architectural options, trade-offs, scalability, reliability, and operational requirements.
- Maintain a high standard for service reliability, performance, observability, and maintainability.
- Proactively identify technical and operational problems and take ownership of resolving them rather than allowing issues to remain unaddressed.
- Use modern AI-assisted development tools thoughtfully while maintaining strong ownership of system design, engineering decisions, and code quality.
- Contribute to a collaborative engineering culture by sharing knowledge, supporting teammates, and helping others solve technical challenges.
Requirements
~2 min read- 5+ years of professional experience in backend engineering, platform engineering, or a closely related discipline.
- Extensive professional experience with Python, the primary programming language used in the environment.
- Hands-on experience developing and operating software on a public cloud platform such as AWS, GCP, or Azure, or working with self-managed Kubernetes; experience with AWS is particularly relevant.
- Strong experience designing and building distributed, high-load services and APIs.
- Solid understanding of data structures, algorithms, and the trade-offs involved in selecting and implementing them.
- Strong system-design mindset, with the ability to design solutions before implementation and understand the architectural implications of technical decisions.
- Experience working effectively with modern AI-assisted coding and development tools, combined with the judgment to understand their appropriate use and limitations.
- High level of ownership, initiative, and accountability, with a proactive approach to identifying and solving problems.
- Strong communication and collaboration skills, with a friendly and supportive approach to working with teammates and cross-functional partners.
- Experience with Rust, C, C++, or Go is a strong advantage.
- Previous experience developing or contributing to ML platforms or machine learning infrastructure is highly valued.
- Experience setting up and operating vector databases such as Qdrant, Milvus, Weaviate, OpenSearch, or pgvector is a plus.
- Experience with model serving or inference infrastructure, including LLM workloads, is advantageous.
- Experience with Infrastructure as Code tools such as Terraform is a plus.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 30, 2026
- First seen
- September 30, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- -1
- Repost count
- 0
- Trust Level
- 80%
- Scored at
- September 30, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.