Job Description
Our client is a leading AI infrastructure company building and operating large-scale compute environments that power next-generation AI workloads. Their teams design, deploy, and operate high-performance data center infrastructure at massive scale, with a focus on speed, reliability, and operational excellence.
They are seeking a Compute Deployment Engineer to own server and accelerator cluster bring-up from facility readiness through production deployment. This role is ideal for someone who thrives in fast-paced environments, takes ownership of large-scale deployments, and enjoys solving complex infrastructure challenges.
Responsibilities:
This Compute Engineer will bring gigawatts of accelerators from first power-on to production. Facility availability to ready-for-service across thousands of racks per site, with a new data hall landing every few weeks.
They will make rack qualification faster than the fleet grows. Firmware baselines, burn-in, and cluster validation proven on every rack before a customer workload touches it, at a pace that never becomes the critical path.
They must be able to scale by tooling, not headcount. Deployed megawatts grow severalfold next year while the team stays near-flat, because anything done twice by hand becomes software.
Must own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
Ability to qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
Ability to travel 20-30% of the time to our Data Centers and Labs, as needed.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
This person has: brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
Worked deep within Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups.
Has automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software.
Has worked physically in data halls, racking, cabling, and swapping components, and you're just as effective acting as remote hands or directing them.
Be able to triage failures methodically across hardware, firmware, and software, isolating the fault to a component before reaching for a fix.
Willing to travel for turn-up windows when a new data hall comes online.
Nice to Have Skills & Experience
Bonus: Kubernetes-based bare-metal provisioning. Accelerator platform bring up (NVIDIA, AMD, or custom). Burn-in and stress harness design. DCIM and inventory tooling.
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.