3h ago
New
CAD 110000-130000/yr

Site Reliability Engineer II (AI Platform)

CanadaCanada·Torontomid
EngineeringDevops Engineer
3 views0 saves0 applied

Quick Summary

Overview

This hybrid role requires working in the office two days per week. With millions of diners, 70,000+ restaurant partners and 25+ years of experience, OpenTable, part of Booking Holdings, Inc.

Technical Tools
EngineeringDevops Engineer

This hybrid role requires working in the office two days per week.

With millions of diners, 70,000+ restaurant partners and 25+ years of experience, OpenTable, part of Booking Holdings, Inc. (NASDAQ: BKNG), is an industry leader with a passion for helping restaurants thrive. Our world-class technology empowers restaurants to focus on what matters most – their team, their guests, and their bottom line – while enabling diners to discover and book the perfect restaurant for every occasion. 

Every employee at OpenTable has a tangible impact on what we do and how we do it. You’ll also be part of a global team and its portfolio of metasearch brands. Hospitality is all about taking care of others, and it defines our culture.

About the Role

~1 min read

As a Site Reliability Engineer II on the Serving Platforms team within Infrastructure Engineering, you will design, automate, and manage the core container stack and infrastructure powering our global business applications. Operating in a high-scale, self-hosted environment, you will serve as a subject matter expert for Kubernetes, Linux systems, and cloud-native automation, directly driving the reliability, security, and efficiency of our platform. In this role, you will collaborate with cross-functional engineering teams worldwide, lead greenfield infrastructure projects, resolve complex incidents, and build self-service capabilities that empower application developers across the organization.

Responsibilities

~1 min read
  • →Maintain, tune, and ensure high availability for the low-level Linux operating system and Kubernetes control plane across our self-hosted bare-metal infrastructure.
  • →Architect, build, and maintain scalable container management, configuration management, and automation tools across global environments.
  • →Investigate, resolve, and conduct root-cause analysis for complex infrastructure disruptions and performance bottlenecks at the system call level.
  • →Participate in high-impact platform engineering projects and collaborate with globally distributed engineering teams to drive infrastructure standardization.
  • →Participate in the team's on-call rotation to support critical production systems and ensure operational resilience.
  • →Develop and maintain self-service tools, automation pipelines, and robust infrastructure monitoring to eliminate manual operational overhead.

Requirements

~1 min read
  • 5+ years of hands-on Linux experience (e.g., Ubuntu, CentOS) with expertise in kernel tuning (sysctl), process management (cgroups/namespaces), system calls, and performance optimization.
  • 3+ years of experience using configuration management systems such as Puppet, Chef, Ansible, or SaltStack in production environments.
  • Proven experience building, operating, and troubleshooting bare-metal Kubernetes clusters from the ground up, including control plane, etcd, and CNI plugin management.
  • Proficiency with continuous system automation and scripting in languages such as Go, Python, Ruby, Perl, or Bash.
  • Demonstrated experience responding to live service disruptions, leading root-cause analysis, and operating messaging systems (e.g., Kafka or RabbitMQ) in production.
  • Experience operating, scaling, and monitoring AI/ML or LLM-powered services and workloads in high-concurrency production environments.
  • Hands-on expertise with public cloud providers (AWS, GCE, or Azure) and containerized CI/CD pipelines (e.g., GitHub, Jenkins, CircleCI, Docker).
  • Experience with distributed key-value stores (e.g., Consul, etcd, Zookeeper, Redis) and enterprise observability/alerting tools (e.g., Prometheus, Sensu).
  • Familiarity with server virtualization infrastructure (e.g., Proxmox, VMware, Xen, OpenStack) and low-level networking concepts (IPtables/NFTables, routing, load balancing).
  • Experience developing and maintaining OS-level software packaging (RPM/DEB) and participating in globally distributed software engineering teams.

This posting is for an existing vacancy.

What We Offer

~2 min read
✓Work from (almost) anywhere for up to 20 days per year
✓Focus on mental health and well-being:
✓Company-paid therapy sessions through SpringHealth
✓Company-paid subscription to Headspace
✓Annual company-wide week off a year - the whole team fully recharges (and returns without a pile-up of work!)
✓Paid parental leave
✓Generous paid vacation + time off for your birthday
✓Paid volunteer time
✓Focus on your career growth:
✓Development Dollars
✓Leadership development
✓Access to thousands of on-demand e-learnings
✓Travel Discounts
✓Employee Resource Groups
✓20 days of paid time off
✓Private health and dental insurance
✓Life and Disability insurance

At OpenTable, we pride ourselves on fostering a global and dynamic work environment. As a team member with us, you will benefit from a schedule tailored to accommodate a global workforce operating across multiple time zones. While the majority of your responsibilities may align with conventional business hours, there will be instances where you are expected to manage communications - via calls, Slack messages, or emails - outside of regular working hours to effectively collaborate with international colleagues, respond to restaurant partners, and/or address urgent matters. OpenTable will always abide by and consider local laws and regulations.

We’re committed to creating a workplace where everyone feels they belong and can thrive. We know the best ideas come when we bring different voices to the table, so we're building a team as dynamic as the diners and restaurants we serve—and fostering a culture where everyone feels welcome to be themselves.

If you need accommodations during the application or interview process, or on the job, we’re here to support you. Please reach out to your recruiter to request any accommodations.

#LI-Hybrid

Location & Eligibility

Where is the job
Toronto, Canada
On-site at the office
Who can apply
Open to applicants worldwide

Listing Details

Posted
September 29, 2026
First seen
September 30, 2026
Last seen
September 30, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
67%
Scored at
September 30, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Site Reliability Engineer II (AI Platform)CAD 110000-130000