fulcrumdigital8d ago
New
New
SRE (Application Support + Dev-Ops + Automation)
EngineeringDevops Engineer
0 views0 saves0 applied
Quick Summary
Overview
Who are we Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation.
Technical Tools
EngineeringDevops Engineer
Who are we Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing. Requirements As part of the Business Operations team, you will: Independently execute key elements of projects/processes within the Site Reliability Engineering area by applying in-depth knowledge of their discipline and area best practices to effectively resolve problems and roadblocks as they occur. Assist in evaluating operational requirements and developing technical solutions within existing frameworks. Support automation and scripting efforts to improve operational workflows and incident response processes. Troubleshoot and resolve routine and some complex system issues, escalating when necessary to maintain system health. Contribute to documentation, knowledge sharing, and best practices to enhance team operational procedures. Collaborate with development teams and stakeholders to ensure reliability solutions align with technical and business needs. Participate in reviews and quality assurance activities to uphold system stability standards. May contribute to solution development for new products/services and/or manage smaller project/initiatives as an experienced individual contributor with specialized knowledge within the Site Reliability Engineering area. Role qualifications: The ideal candidate will apply the following skills independently in routine and moderately complex situations, requiring occasional guidance typically only in unfamiliar or highly complex scenarios. They will demonstrate growing consistency and reliability in applying the skills. Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement. Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar. Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability. Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency. Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations. Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations. Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage. IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs. Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability.
Location & Eligibility
Where is the job
Dublin, Ireland
On-site at the office
Listing Details
- Posted
- September 4, 2026
- First seen
- September 11, 2026
- Last seen
- September 11, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 29%
- Scored at
- September 11, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application · ~5 min on fulcrumdigital's site
Please let fulcrumdigital know you found this job on Jobera.
3 other jobs at fulcrumdigital
View all →Explore open roles at fulcrumdigital.
Browse Similar Jobs
Security2.3kEngineering Manager2.3kFullstack Developer2.3kDevOps & Infrastructure2.2kSoftware Architect2kQa Engineer1.8kBackend Developer1.6kMechanical Engineer1.5kSecurity Engineer1.5kElectrical Engineer1.2kFrontend Developer1.2kMobile Developer1.1kProject Engineer1kData Engineering967Backend Engineering941Design Engineer885Product Engineer551Embedded Engineer546Automation Engineer540Process Engineer507
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.