MLOps / AI Operations Engineer
Quick Summary
Job Description & Summary The opportunity Industrialize AI delivery through automated deployment, evaluation operations, observability, reliability engineering and transparent consumption management.
Industrialize AI delivery through automated deployment, evaluation operations, observability, reliability engineering and transparent consumption management.
Responsibilities
~1 min read· Build CI/CD pipelines for AI services, prompts, agent configurations, infrastructure and evaluation assets.
· Automate environment provisioning, testing, deployment, rollback and release evidence.
· Implement tracing, logging, model and agent monitoring, alerts and operational dashboards.
· Operationalize evaluation thresholds, incident handling and continuous-improvement loops.
· Monitor latency, capacity, token usage, infrastructure consumption and cost drivers.
· Define runbooks, service ownership and production support handover.
Requirements
~1 min read· 4+ years in DevOps, platform engineering, ML engineering, SRE or cloud operations.
· Strong automation, containers, cloud services, observability and Infrastructure as Code capability.
· Experience deploying or operating ML, generative AI or distributed application workloads.
· Understanding of release controls, reliability, security and cost optimization.
· Hands-on experience with GitHub Actions, Azure DevOps, GitLab CI or equivalent, plus Infrastructure as Code using Terraform, Bicep or comparable tooling.
· Strong container and orchestration capability using Docker and Kubernetes, together with experience deploying AI or agent services across cloud and hybrid environments.
· Experience operating model and prompt assets, agent configurations, evaluation datasets and release evidence using MLflow, platform-native registries or equivalent lifecycle tooling.
· Practical implementation of agent tracing and observability using OpenTelemetry and tools such as LangSmith, MLflow, Langfuse, Azure Monitor, Prometheus or Grafana.
· Ability to monitor model and agent quality, tool failures, retrieval performance, latency, token usage, cost, capacity and workflow-level service indicators.
· Experience with progressive delivery, rollback, secrets management, vulnerability scanning, incident response and reliability practices for non-deterministic AI systems.
· Deployment frequency and success rate
· Mean time to detect and restore
· Evaluation and monitoring coverage
· Service reliability and latency
· Cost and consumption transparency
· Other members of the AI Transformation & Agentic Systems Practice
· PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists
· Client business owners, product owners, technology teams and operational users
· Technology alliance and implementation partners where relevant
· Support proposals, client workshops and market development appropriate to seniority.
· Contribute reusable methods, patterns, code, assets and lessons learned.
· Coach colleagues and participate in the capability’s continuous learning agenda.
· Uphold PwC quality, independence, confidentiality and risk-management requirements.
#LI-BS1 #LI-Hybrid
Location & Eligibility
Listing Details
- Posted
- September 23, 2026
- First seen
- October 3, 2026
- Last seen
- October 6, 2026
Posting Health
- Days active
- 2
- Repost count
- 0
- Trust Level
- 19%
- Scored at
- October 6, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.