Cloud Operations TM Lead
Quick Summary
Bachelor’s degree or an equivalent combination of education and relevant professional experience. 5+ years of experience in cloud operations, infrastructure engineering, DevOps,
This role offers the opportunity to operate and enhance cloud infrastructure supporting large-scale transportation management solutions. You will work closely with solution architects, engineering teams, and business stakeholders to design, deploy, and maintain reliable production environments. The position combines AWS and Kubernetes operations with infrastructure as code, automation, incident management, and full-stack troubleshooting. You will help ensure high availability and operational performance for customer-facing SaaS applications while continuously improving service levels. The role provides broad exposure across cloud platforms, Java and .NET applications, containers, operating systems, networking, and observability. It is ideal for an experienced cloud operations professional who enjoys solving complex production challenges and building more efficient, automated operational processes.
-
Manage and support SaaS Kubernetes applications running in AWS, ensuring reliable and efficient production operations.
-
Monitor event-driven alerts, investigate incidents, and resolve operational issues to maintain application availability and service-level objectives.
-
Triage and resolve tickets raised by client services teams, collaborating with relevant technical stakeholders when escalation is required.
-
Work with solution architects and business stakeholders to identify, design, develop, and deploy infrastructure engineering solutions.
-
Support the rapid and controlled rollout of new technologies and capabilities across cloud environments.
-
Troubleshoot operating system, application, virtualization, and containerization issues across the technology stack.
-
Support large-scale Java applications and contribute to troubleshooting .NET-based applications where required.
-
Implement and maintain Infrastructure as Code using Terraform and related technologies.
-
Develop automation and operational tooling using Jenkins, Ansible, Python, Terraform, and similar technologies.
-
Collect and analyze system and application performance data, identify root causes, and implement appropriate tuning and optimization.
-
Manage Kubernetes environments, including Helm-templated deployments and containerized applications such as Tomcat.
-
Perform full-stack troubleshooting across cloud infrastructure, applications, containers, networking, and supporting services.
-
Use Git and established source-control practices to manage automation and infrastructure code.
-
Contribute to the continuous improvement of operational processes, automation, reliability, and service-level performance.
Requirements
~2 min read-
Bachelor’s degree or an equivalent combination of education and relevant professional experience.
-
5+ years of experience in cloud operations, infrastructure engineering, DevOps, or software development environments.
-
3+ years of experience supporting large-scale Java applications.
-
3+ years of hands-on experience supporting AWS-based environments.
-
Extensive practical experience with Infrastructure as Code principles and technologies, particularly Terraform.
-
Strong understanding of Linux operating systems and cloud-based production environments.
-
Hands-on experience implementing, managing, and troubleshooting Kubernetes solutions.
-
Broad automation experience using technologies such as Jenkins, Terraform, Ansible, and Python.
-
Experience collecting performance metrics, analyzing system behavior, troubleshooting issues, and tuning applications or infrastructure.
-
Strong knowledge of application containers, particularly Tomcat.
-
Experience with full-stack troubleshooting and Git/source-control management.
-
Multi-cloud experience, particularly with AWS and Azure, is advantageous.
-
Experience supporting .NET applications and working with Dynatrace is a plus.
-
Familiarity with Helm-based Kubernetes environments and microservices architectures is desirable.
-
Understanding of common application protocols and messaging technologies, including TCP/IP, HTTP, SOAP, SMTP, REST APIs, XML/JSON, JDBC, and JMS/MQ.
-
Knowledge of application threading and concurrency concepts, as well as troubleshooting response-time, connectivity, authentication, authorization, and configuration issues.
-
Strong analytical, problem-solving, and communication skills with the ability to work effectively across engineering, architecture, business, and customer-facing teams.
-
Ability to work independently in a production-focused environment and continuously improve operational processes.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 30, 2026
- First seen
- September 30, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 30, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.