NuevoEngineering

Director, Sustaining Engineering - Spark

Crusoe

Denver3h ago

Aplica ahora

Acerca del rol

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role

Crusoe is a vertically integrated AI Factory company with a mission to accelerate the abundance of energy and intelligence. Our competitive advantage, "Speed is the only moat," is directly tied to our ability to rapidly design, manufacture, and deploy our own modular power and compute infrastructure.

We are seeking a Sustaining Engineering Leader to own how Spark performs after it is deployed. Product Management defines the next generation. New Product Introduction delivers it into production. You own each generation from the moment it leaves ramp: fleet reliability, field failures, configuration, obsolescence, and the design changes that keep deployed units performing for their full service life.

You do not set the roadmap and you do not run the factory. You own the installed base, and you are the evidence source Product Management and NPI both depend on. You will start as the single-threaded owner of this function and build the team as the fleet grows.

First Year Mandate

Spark deployments begin in early 2027. You will build the system before there is a fleet to run it on.

  • Stand up FRACAS, the root cause evidence standard, and the Failure Review Board.
  • Establish the as-built configuration baseline before there are units to reconcile.
  • Define spares, service intervals, and repair versus replace policy for the first generation.
  • Participate in the production release and ramp readiness reviews for the first deploying generation as the receiving organization.

I. Fleet Ownership & Field Issues

  • Transfer of Ownership: Take technical ownership of each generation once it completes production ramp, through a defined handover from New Product Introduction covering the released configuration, open issues, and known risks.
  • Fleet Reliability: Own failure rate by subsystem, mean time between failures, and repeat failure rate, and define the telemetry every unit must report to make them measurable. Data Center Facility Operations owns uptime and time to repair.
  • Closed Loop Corrective Action: Run FRACAS and chair the Failure Review Board. Close every failure with a root cause, a verified fix, and proof it worked. A unit that fails while built within released tolerances is a design issue; attribution unresolved after ten business days proceeds as a Spark-funded fix and escalates to the Spark GM.
  • Containment: For safety or fleet-wide risk, contain first and settle attribution after. Own the retrofit decision and its sequencing.

II. Change, Configuration & Obsolescence

  • Post-Release Change Control: Own engineering change for generations in the field, with cost and fleet impact quantified before approval. New Product Introduction owns change control up to production release; you own it after.
  • Configuration Control: Own the configuration baseline and define what Operations records. Reconcile as-maintained against baseline so a change targets exactly the units that need it.
  • Obsolescence: Run a proactive DMSMS program. Model last time buys against fleet demand and qualify alternates before a shortage forces the choice.
  • Retrofit & Maintenance: Issue engineering change packages, kit definitions, preventive maintenance intervals, and acceptance criteria. Operations executes and returns the records; you verify effectiveness.

III. Feeding Product Management and NPI

  • Requirements into Product:

Convert fleet evidence into quantified design requirements. Product Management decides what enters the PRD.

  • Qualification Requirements into NPI:

Supply field-derived requirements and acceptance criteria, so a failure the fleet has already seen must be designed out and proven before the next generation is released.

  • Total Cost of Ownership: Own the reliability and service inputs to Spark unit economics.
  • Serviceability Advocacy: Represent serviceability and maintainability in design reviews. You hold no approval authority; you make the case with fleet data.

Education

  • Required: Bachelor's degree in Mechanical, Electrical, Reliability, or a closely related engineering discipline.

Experience

  • 10 to 15 years in sustaining, reliability, or product support engineering for deployed capital equipment or infrastructure hardware.
  • Direct ownership of a fielded installed base, with field data converted into design change that measurably reduced failure rate or service cost.
  • Obsolescence and lifecycle management on a product whose service life exceeds that of the components inside it.
  • Experience standing a function up from nothing, including the processes and the partner agreements behind it.

Technical Depth

  • Reliability: FRACAS, structured root cause analysis, and mean time between failures and population failure rate analysis.
  • Change & Configuration: PLM, change order workflow, as-built control, and serial level effectivity.
  • Obsolescence: DMSMS practice, end of life monitoring, last time buy modeling, and alternate part qualification.
  • Systems Fluency: Electrical, mechanical, and thermal command sufficient to adjudicate root cause on a modular power and compute product, including liquid cooling.
  • Service Economics: Spares, service cost per unit, and the link between availability and revenue.

Additional Qualifications

  • Experience with modular, containerized, or prefabricated infrastructure products in the field.
  • Familiarity with liquid cooled data center hardware, including CDU and rack level cooling interfaces.
  • Mountaineer Spirit: The persistence to chase a root cause past the easy answer, and the judgment to know when a fleet-wide fix is worth its disruption.

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowance
  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $225,000-$255,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Fuente: la propia página de carreras del empleador.

Empleos similares