Jobs / United States / Nvidia Corporation

Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

Nvidia Corporation · 🇺🇸 US, CA, Santa Clara

Sponsorship verdict

Sponsorship possible

One solid signal, not two — worth applying, and worth asking about sponsorship early.

  • Employer is on a government sponsor recordThe US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
  • The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
  • No salary bar for this routeH-1B has no fixed salary bar: the employer must pay at least the prevailing wage for the role and area. Cap-subject employers enter a lottery weighted by wage level. Source: https://www.federalregister.gov/documents/2025/12/29/2025-23853/weighted-selection-process-for-registrants-and-petitioners-seeking-to-file-cap-subject-h-1b, rules effective 2026-02-27.
  • What Nvidia Corporation paid sponsored hires in similar roles701 certified filings for “Engineer Senior Systems Software” (Software Developers) in CA: $173k–$214k, median $190k. Most were filed at wage level IV (82%) — 4 lottery entries, ≈61% projected selection odds for cap-subject employers. Source: US Department of Labor LCA disclosure data (Oct 2025 – Jun 2026).
  • Confirmed live todayWhen a source last listed this job as open.

US H-1B: cap-subject employers enter a lottery weighted by wage level — Level I gets 1 entry, Level IV gets 4 (DHS projected selection odds ≈15% at Level I to ≈61% at Level IV). Universities and non-profit research employers are cap-exempt. The $100,000 fee for new petitions from abroad is currently blocked by a court order (appeal pending).

A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.

Start free →

Or apply yourself on the official page →

Sponsor Radar — Nvidia Corporation

2,374 H-1B filings certified since Oct 2025

The US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).

Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Nvidia Corporation →

About the role

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for fleet-scale reliability, working directly with engineering teams of key CSP / hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production. In this role, you will augment NVIDIA's internal software/firmware and quality teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of reliability software/firmware architecture, methodology, incorporate their fleet telemetry and failure data into NVIDIA's improvement priorities, and validate that reliability improvements measured in the lab translate to real customer environments. Your cross-CSP visibility enables you to distinguish systemic architectural gaps from environmental or configuration-specific issues that no single customer engagement could identify alone. What you'll be doing: • Drive reliability work streams with CSP engineering teams — ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architecture • Gather and synthesize CSP fleet reliability data — identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teams • Define consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practices • Conduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specific • Drive fleet health monitoring integration architecture — ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automation • Define burn-in reliability test environment and cluster certification criteria in collaboration with quality teams, validating with customers that criteria are meaningful • Collaborate with CSPs to ensure reliability-related integration work (health monitoring deployment, telemetry pipeline, alerting configuration) is complete ahead of at-scale launch • Develop predictive failure models using fleet telemetry and validate their effectiveness in customer environments What we need to see: • 15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges. • BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience) • Deep expertise in multi-NUMA, rack-scale system software and firmware. Statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classification • Experience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlation • Understanding of hardware failure modes in large-scale GPU/accelerator deployments — ability to classify and prioritize across compute, interconnect, memory, power, and thermal domains • Experience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems. Familiarity with predictive maintenance or anomaly detection approaches applied to fleet health data • Customer obsession — genuine passion for understanding fleet reliability challenges at scale and translating them into actionable engineering priorities • Strong communication — ability to present statistical reliability findings to both deep technical audiences and executive leadership. Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authority Ways to stand out from the crowd: • Experience in flee

View the official posting →

Source: Employer career site (Workday) First seen: 2026-08-30 Last confirmed: 2026-10-02 How our data works → Report this job

Similar opportunities