Quick Summary
caching strategies, model routing,
Coursera and Udemy are now one company, bringing together two mission-driven brands to create the world’s most powerful platform for turning learning into progress. Together, we help more than 300 million learners and 12,000+ enterprise customers build the skills they need for a world being reshaped by AI. Read more about the combined company by visiting our blog.
What We Offer
~2 min readAI is transforming how people learn, work, and grow, and the need for new skills has never been greater. Coursera brings trusted content and credentials from leading university and industry partners, while Udemy brings a dynamic skills marketplace and global network of real-world experts. By combining these strengths, we can connect more people and organizations to the skills they need, when they need them.
By joining our team, you’ll have the opportunity to reshape how the world learns and applies skills—and help millions of people participate in the new economy. Bring your ideas, expertise, and perspective to meaningful work that can make a difference at global scale.
As an AI Platform and LLMOps Engineer, you will join a fast-paced innovation team at Coursera working on AI enablement for our organization and customers. As an internal focus area, this team builds and maintains the infrastructure & tooling required to drive meaningful AI adoption across the organization, providing a platform for non-technical team members to build and deploy tools and transforming entire business functions into AI-native operating models. As a customer-facing focus area, this team builds custom AI-powered solutions tailored for our enterprise customers supporting them in their journey of AI adoption and upskilling beyond just the content on our platforms.
You will own the operational backbone that lets our agentic AI systems run reliably in production — from prompt and pipeline versioning to evaluation, monitoring, incident and cost management. You will work closely with a cross-functional team of Software Engineers, AI Specialists and Product Managers – helping turn promising AI prototypes into reliable, observable, and cost-effective production systems.
Responsibilities
~1 min read- →Build and maintain the operational tooling for our LLM-powered systems, including prompt/pipeline versioning, evaluation harnesses, and CI/CD workflows tailored to non-deterministic AI outputs
- →Implement monitoring and observability for production LLM systems — tracking latency, token usage, cost per request, output quality, and drift over time
- →Design and run automated evaluation suites to catch regressions, hallucinations, and quality degradation before they reach customers
- →Manage RAG pipelines and vector store infrastructure, keeping retrieval sources fresh, accurate, and performant
- →Implement safety and compliance guardrails — content filtering, PII redaction, and access controls — in line with enterprise data privacy and residency requirements
- →Own cost governance for LLM usage: caching strategies, model routing, and usage reporting to keep spend predictable as adoption scales
- →Collaborate closely with Product Managers and Senior Engineers to scope operational requirements and translate them into reliable systems
- →Participate in sprint planning, technical grooming, and retrospective discussions
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering or a related field
Requirements
~2 min read- 4+ years of experience in a software engineering role, with a solid background in Backend Engineering and DevOps fundamentals
- Hands-on experience operating or supporting LLM-powered systems in production (via APIs, RAG pipelines, frontier/open weight/fine-tuned models)
- Proficiency in a backend language such as Python, Java, TypeScript or Go, and familiarity with Docker and container orchestration (Kubernetes)
- Experience implementing APIs, working with SQL and NoSQL databases, and writing automated tests
- Working knowledge of at least one cloud platform (AWS, GCP, or Azure), across both managed and self-hosted services
- Familiarity with infrastructure-as-code (e.g., Terraform) and CI/CD tooling (e.g., GitHub Actions, Jenkins)
- Comfort debugging production issues involving non-deterministic systems, with attention to detail and a bias towards reliability
- Strong belief in engineering quality and building tooling that creates leverage for others
- Experience with LLMOps-specific tooling such as LangFuse, LangSmith, Weights & Biases-style evaluation frameworks, or vector databases like Pinecone, Weaviate, pgvector
- Working knowledge of RAG architectures, MCP implementation & governance, agent orchestration frameworks like LangGraph and LLM Gateways such as LiteLLM, Open Router & enterprise AI ecosystems such as Vertex AI, Bedrock
- Exposure to prompt management and versioning practices treated as code along with model access management, deterministic guardrails and evals
- Understanding of AI governance considerations — data privacy, residency, and compliance in enterprise AI deployments
- Ability to work in a fast-paced, ambiguous environment with a proactive, ownership-driven mindset
- Strong communication skills and comfort collaborating across engineering, product, and strategy functions
Location & Eligibility
Listing Details
- Posted
- October 6, 2026
- First seen
- October 6, 2026
- Last seen
- October 6, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 57%
- Scored at
- October 7, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.