Senior DevOps Engineer
Quick Summary
5+ years of experience in DevOps or Site Reliability Engineering roles. Proven experience supporting production systems with high availability and reliability requirements.
This role focuses on building and operating reliable, scalable cloud infrastructure for a fast-growing AdTech and e-commerce platform. You’ll take ownership of production infrastructure, containerized applications, deployment workflows, monitoring, and operational reliability. Working primarily with AWS and Kubernetes, you’ll help optimize systems for performance, scalability, availability, and cost efficiency. You’ll partner closely with development teams to improve CI/CD processes and make deployments smoother and more reliable. The role also offers opportunities to strengthen infrastructure automation, observability, and Infrastructure-as-Code practices. You’ll work in a flexible, remote environment where proactive problem-solving and strong ownership are highly valued.
- Manage, deploy, and maintain containerized applications using Kubernetes in production environments.
- Manage AWS infrastructure across services including EC2, VPC, S3, IAM, RDS, and OpenSearch, ensuring scalability, reliability, security, and cost efficiency.
- Monitor infrastructure and application performance, maintaining and improving alerting and observability systems to support high availability.
- Collaborate closely with development teams to streamline CI/CD pipelines and optimize deployment workflows.
- Troubleshoot infrastructure and application issues, investigate root causes, and drive timely incident resolution.
- Continuously improve infrastructure automation, operational processes, and DevOps best practices.
- Support performance tuning and cost optimization of EC2 instances and other cloud resources.
- Contribute to GitOps and Infrastructure-as-Code practices where appropriate.
- Build and maintain dashboards and monitoring solutions using tools such as Prometheus, Grafana, and AWS CloudWatch.
Requirements
~1 min read- 5+ years of experience in DevOps or Site Reliability Engineering roles.
- Proven experience supporting production systems with high availability and reliability requirements.
- Strong hands-on experience with Kubernetes for container orchestration and production deployments.
- Solid understanding of AWS services, particularly EC2, VPC, S3, and IAM.
- Practical experience managing EC2 instances, including performance tuning and cloud cost optimization.
- Experience with monitoring and observability tools such as Prometheus, Grafana, and AWS CloudWatch.
- Strong scripting skills with Python and/or Bash for automation and operational tasks.
- Strong troubleshooting and problem-solving abilities, with a proactive approach to infrastructure reliability.
- Ability to take ownership of infrastructure initiatives while working effectively both independently and within a development team.
- Familiarity with GitOps practices for continuous delivery and infrastructure automation is a plus.
- Experience with Terraform or other Infrastructure-as-Code tools is advantageous.
- Advanced Grafana experience, including custom dashboards and visualizations, is beneficial.
- Experience managing AWS RDS instances is a plus.
- Familiarity with the Elastic Stack, including Elasticsearch, Kibana, and Elastic Cloud, is advantageous.
- Experience with AWS OpenSearch for search and analytics use cases is a plus.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- October 1, 2026
- First seen
- October 1, 2026
- Last seen
- October 4, 2026
Posting Health
- Days active
- 2
- Repost count
- 1
- Trust Level
- 62%
- Scored at
- October 4, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.