Senior Site Reliability Engineer for Fuse team
Quick Summary
Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey.
- We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses.
- We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey.
- We're designing the future of autonomous marketing, taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do.
Join the Fuse team — the team responsible for the item data management capabilities that connect Bloomreach Data Hub with Marketing, Search, Recommendations, and emerging Loomi agent use cases.
Fuse owns and evolves the systems behind Data Hub item collections: ingesting items data, transforming and validating it, managing data schemas and lifecycle, and distributing data reliably to downstream Bloomreach products. Item collections provide a unified source of data that can be used across all Bloomreach products.
Our current areas of focus include:
- Unified items data pipelines: processing records into structured items and keeping data synchronized with Marketing and Search destinations.
- Catalog APIs and lifecycle management: customer-facing and internal APIs, catalog creation and naming, schemas, destinations, migrations, and backward-compatible evolution.
- Scalable storage and indexing: operating and improving systems built on PostgreSQL, Bigtable, Elasticsearch, Solr.
- Reliable jobs execution: submission, queueing, execution, progress reporting, retries, cancellation, rate limiting, and operational tooling.
- Cross-product capabilities: catalog data triggers, multi-dimensional data support, custom item types, catalog data enrichment, recommendations, and semantic catalog profiles for agentic use cases.
As a Senior SRE, you will be the team’s reliability and operability leader. You will work alongside backend engineers, embedded QA, Product, and Engineering Management to make complex product-data systems observable, scalable, safe to release, and straightforward to operate.
Fuse embraces AI-assisted engineering. We expect engineers to use modern coding agents thoughtfully to accelerate investigation, development, testing, documentation, and operational work while retaining full ownership of correctness, security, and production outcomes.
Working from one of our Central European offices (Bratislava, Prague, or Brno), or remotely (Czechia, Slovakia) on a full-time basis, you’ll become a core part of the Engineering organization.
As a P3 Senior SRE at Bloomreach, you are an independent reliability professional who can turn ambiguous operational problems into measurable improvements and lead initiatives end-to-end with minimal day-to-day guidance.
Your challenge will be to make Fuse’s distributed data platform dependable across the complete data path:
Responsibilities
~1 min readIf this position doesn't suit you, but you know someone who might be a great fit, share it - we will be very grateful!
Any unsolicited resumes/candidate profiles submitted through our website or to personal email accounts of employees of Bloomreach are considered property of Bloomreach and are not subject to payment of agency fees.
#LI-Remote
- Own and improve the reliability posture of Fuse services, workers, APIs, queues, storage systems, and destination synchronization pipelines.
- Establish meaningful SLIs, SLOs, and error budgets for customer-facing APIs, asynchronous jobs, catalog data freshness, destination synchronization, and indexing.
- Build end-to-end observability across Data Hub item collections, from API request and job submission through processing, persistence, indexing, and downstream delivery.
- Ensure engineers can trace a workspace, item collection, catalog, or job across services without manually correlating disconnected logs and database records.
- Create and maintain actionable dashboards, alerts, and service health views using Grafana, Prometheus-compatible metrics, OpenTelemetry, PagerDuty, and GCP tooling.
- Detect missing, stalled, duplicated, or inconsistent processing before customers or downstream teams report it.
- Improve capacity planning and autoscaling using workload telemetry, queue depth, processing throughput, latency, memory usage, storage growth, and customer-level traffic patterns.
- Reduce noisy alerts and replace symptom-based monitoring with signals tied to customer impact.
- Improve the availability, scalability, and operability of catalog data across PostgreSQL/Cloud SQL, Bigtable, Elasticsearch, GCS, Kafka, and related storage systems.
- Support catalog placement, routing, index lifecycle, shard management, safe migration, and recovery across multiple Elasticsearch clusters.
- Develop safeguards for full replacements, delta updates, deletions, schema changes, destination changes, and catalog reindexing.
- Define and automate data-consistency checks between source records, transformed items, job state, Bigtable, Elasticsearch, and downstream destinations.
- Help establish practical platform limits and quotas for catalog size, API traffic, job concurrency, queue depth, payload size, and expensive operations.
- Partner with engineers on performance testing for large catalogs and high-throughput customer workloads.
- Own and evolve Kubernetes configuration and operational infrastructure for Fuse components.
- Improve deployment automation, progressive rollout, rollback, and validation across development and production environments.
- Make coordinated releases safer when changes span
app/app, Fuse workers, Kubernetes configuration, and PostgreSQL migrations. - Automate operational procedures that currently depend on manual commands, one-off scripts, or specialist knowledge.
- Maintain CI/CD pipelines with tests, linters, dependency management, security checks, image publication, and release verification.
- Create reusable tooling for local development, ephemeral environments, end-to-end testing, load testing, and production diagnosis.
- Ensure runbooks remain executable and are validated through exercises rather than existing only as documentation.
- Participate in and help improve the Fuse L3/on-call rotation.
- Lead incident investigation, mitigation, stakeholder communication, and follow-up for Fuse-owned systems.
- Use logs, metrics, traces, database state, queue state, and Kubernetes signals to diagnose failures across distributed workflows.
- Build safe operational tools for common support activities such as job tracing, queue inspection, rate-limit diagnosis, catalog health checks, and index recovery.
- Facilitate blameless incident reviews and ensure resulting actions address root causes rather than only immediate symptoms.
- Improve the handoff between customer support, L2, Fuse L3, Infrastructure, and dependent engineering teams.
- Reduce recurring support demand by turning incident knowledge into safeguards, automation, tests, dashboards, and clear documentation.
- Help Fuse meet Bloomreach security and compliance requirements, including ISO and SOC 2 controls.
- Enforce least-privilege access, workload identity, service-level authentication and authorization, secret rotation, encryption, and auditability.
- Protect customer isolation across workspaces, item collections, projects, accounts, databases, indexes, buckets, and asynchronous jobs.
- Ensure operational tooling and incident procedures respect production-access restrictions and PII-handling requirements.
- Partner with engineering teams to make security controls observable and testable rather than relying on undocumented assumptions.
- Participate early in the design of new Fuse capabilities so reliability, recovery, observability, limits, and operational ownership are defined before implementation.
- Review designs for failure modes, retry behavior, idempotency, backpressure, ordering, consistency, timeout handling, cancellation, and safe rollout.
- Clarify ownership boundaries and service contracts with teams including Campaigns, Data Pipeline, Integrations, Discovery, Recommendations, Infrastructure, Frontend, and QA.
- Help teams choose architectures that balance immediate delivery with long-term operability and cost.
- Coach engineers in production readiness, operational testing, debugging, and sustainable on-call practices.
Location & Eligibility
Listing Details
- Posted
- April 22, 2026
- First seen
- April 22, 2026
- Last seen
- August 20, 2026
Posting Health
- Days active
- 120
- Repost count
- 0
- Trust Level
- 31%
- Scored at
- August 20, 2026
Signal breakdown

Bloomreach is a cloud-based e-commerce experience platform specializing in marketing automation, product discovery, and content management systems, using AI to personalize customer experiences.
View company profilePlease let Bloomreach know you found this job on Jobera.
4 other jobs at Bloomreach
View all →Explore open roles at Bloomreach.
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.