crusoe
crusoe1d ago
New
USD 140000-165000/yr

Software Engineer II (DCIE)

United StatesUnited States·San Franciscofull-timemid
Software EngineerSoftware Engineering
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.

Technical Tools
Software EngineerSoftware Engineering

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role

~1 min read

We are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The role focuses on the developing and implementing advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters.

The ideal new team member will be a hands-on problem solver who is comfortable working independently. The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet.

Responsibilities

~1 min read
  • Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.

  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.

  • Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.

  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.

  • Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.

  • Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.

  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems

  • 2-3 years of software engineering experience.

  • The ability to identify a problem, rapidly develop a scalable solution and ship it.

  • Ability to lean in and assist team members working on critical or complex technical initiatives.

  • Ability to set the technical direction for a specific project and execute.

  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)

  • Strength in at least one programming language - Go, Python, Java, Rust.

  • Strong analytical and problem-solving skills.

  • Excellent communication and collaboration skills.

  • Ability to work independently and within a team

Nice to Have

~1 min read
  • Experience with Temporal and Kubernetes.

  • Experience working directly with hardware vendors.

  • Background in large-scale GPU fleet operations or hyperscale data center environments.

What We Offer

~1 min read
Industry competitive pay
Restricted Stock Units in a fast-growing, well-funded technology company
Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
Employer contributions to HSA accounts
Paid Parental Leave
Paid life insurance, short-term and long-term disability
Teladoc
401(k) with a 100% match up to 4% of salary
Generous paid time off and holiday schedule
Cell phone reimbursement
Tuition reimbursement
Subscription to the Calm app
MetLife Legal
Company paid commuter benefit; $50 per pay period

Location & Eligibility

Where is the job
San Francisco, United States
On-site at the office
Who can apply
US

Listing Details

Posted
August 4, 2026
First seen
August 4, 2026
Last seen
August 4, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
52%
Scored at
August 4, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

crusoeSoftware Engineer II (DCIE)USD 140000-165000