Cloud Operations WM Lead
Quick Summary
Bachelor’s degree or equivalent combination of education and relevant professional experience. 5+ years of experience in cloud operations, infrastructure engineering, DevOps,
This role is an opportunity to support and evolve cloud infrastructure powering large-scale warehouse management solutions. You will work closely with solution architects, engineering teams, and business stakeholders to design, deploy, and operate reliable cloud environments. The position combines hands-on AWS and Kubernetes operations with automation, infrastructure as code, incident response, and full-stack troubleshooting. You will play a key role in maintaining production customer environments and improving operational reliability and service levels. The environment is technically diverse, spanning Java applications, containers, operating systems, virtualization, and cloud infrastructure. It is well suited to an experienced cloud operations professional who enjoys solving complex production challenges and continuously improving how systems are operated.
-
Manage SaaS Kubernetes applications running in AWS and support reliable operation of production customer environments.
-
Monitor event-driven alerts, investigate incidents, and troubleshoot issues to minimize service disruption and maintain operational SLAs.
-
Triage and resolve tickets raised by client services and other internal stakeholders.
-
Work with solution architects and business stakeholders to identify, design, develop, and deploy infrastructure engineering solutions.
-
Execute plans for introducing new technologies and capabilities into cloud environments.
-
Troubleshoot operating system, application, virtualization, containerization, and infrastructure issues across the technology stack.
-
Support large-scale Java applications and diagnose application and environment-related performance or reliability issues.
-
Implement and maintain Infrastructure as Code using technologies such as Terraform.
-
Develop automation and operational processes using tools such as Ansible, Python, and Azure DevOps to improve efficiency and consistency.
-
Collect and analyze performance data, identify root causes, and tune infrastructure and applications accordingly.
-
Manage and troubleshoot Kubernetes environments, including containerized applications and Helm-based deployments.
-
Work with application technologies such as IIS and Tomcat and provide full-stack troubleshooting across infrastructure and application layers.
-
Use Git and other source-control practices to manage operational and automation code.
-
Support geographically distributed environments and contribute to reliable operations across hosted cloud and data-center infrastructure.
Requirements
~2 min read-
Bachelor’s degree or equivalent combination of education and relevant professional experience.
-
5+ years of experience in cloud operations, infrastructure engineering, DevOps, or software development environments.
-
3+ years of experience supporting large-scale Java applications.
-
3+ years of hands-on experience supporting AWS-based production environments.
-
Experience working with OCI-based environments, with multi-cloud exposure considered an advantage.
-
Strong hands-on experience with Kubernetes administration, implementation, and troubleshooting.
-
Extensive knowledge of Infrastructure as Code principles and practical experience with Terraform.
-
Strong automation skills using tools such as Python, Ansible, Azure DevOps, Terraform, or comparable technologies.
-
Working knowledge of both Windows and Linux operating systems.
-
Experience supporting application containers and technologies such as IIS and Tomcat.
-
Strong troubleshooting and analytical skills, including experience collecting performance data, diagnosing issues, and tuning systems.
-
Experience with Git or other source-control management systems.
-
Understanding of microservices architectures and common application protocols such as TCP/IP, HTTP, SOAP, SMTP, REST APIs, XML/JSON, JDBC, and JMS/MQ is advantageous.
-
Experience with Dynatrace, Helm-templated Kubernetes environments, Java application management, or application threading and concurrency is a plus.
-
Ability to investigate issues involving response times, connectivity, authentication, authorization, configuration, and other environment-related factors.
-
Strong communication and collaboration skills, with the ability to work effectively with technical teams, architects, business stakeholders, and customer-facing teams.
-
Ability to work independently, respond effectively to incidents, and continuously improve operational processes in a fast-paced environment.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 30, 2026
- First seen
- September 30, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 30, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.