Jobs / United States / Nvidia Corporation
Technical Product Manager - AI Infra Resilience
Nvidia Corporation · 🇺🇸 3 Locations
Sponsorship verdict
Sponsorship possible
One solid signal, not two — worth applying, and worth asking about sponsorship early.
- Employer is on a government sponsor recordThe US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
- The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
- No salary bar for this routeH-1B has no fixed salary bar: the employer must pay at least the prevailing wage for the role and area. Cap-subject employers enter a lottery weighted by wage level. Source: https://www.federalregister.gov/documents/2025/12/29/2025-23853/weighted-selection-process-for-registrants-and-petitioners-seeking-to-file-cap-subject-h-1b, rules effective 2026-02-27.
- What Nvidia Corporation paid sponsored hires in similar roles7 certified filings for “Product Manager” (Marketing Managers) in CA: $180k–$217k, median $201k. Source: US Department of Labor LCA disclosure data (Oct 2025 – Jun 2026).
- Confirmed live todayWhen a source last listed this job as open.
US H-1B: cap-subject employers enter a lottery weighted by wage level — Level I gets 1 entry, Level IV gets 4 (DHS projected selection odds ≈15% at Level I to ≈61% at Level IV). Universities and non-profit research employers are cap-exempt. The $100,000 fee for new petitions from abroad is currently blocked by a court order (appeal pending).
A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.
Or apply yourself on the official page →
Sponsor Radar — Nvidia Corporation
The US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Nvidia Corporation →
About the role
GPU clusters fail. Workloads crash at 3am, AI teams file tickets pointing at hardware, hardware teams point back at software, and the actual cause stays unknown until someone with deep enough context digs through DCGM metrics, XID history, and NCCL traces. That forensics work shouldn't fall to a human on-call every time. We're building the platform that automates it, and we need a PM who has been that on-call engineer. We're hiring a Senior Product Manager to own workload failure attribution: how a GPU cluster figures out whether a failed job was caused by hardware, system software, or application code, and how that verdict gets surfaced to operators, schedulers, and the open source developer surface built on top of it. You will own the roadmap and work directly with the engineering and operations communities who deploy and extend this platform. This role is part of NVIDIA's DSX Platform, NVIDIA's open, modular infrastructure software for designing, operating, and optimizing AI factories at scale. No prior product management title required. If you are a deeply technical product manager, architect, SRE, or infra engineer who has lived the GPU cluster failure-triage problem, we want to hear from you. What You'll Be Doing: • Drive the product vision, roadmap, and delivery for workload failure attribution, in close collaboration with engineering on architecture and prioritization. • Identify gaps in how customers and partners triage GPU cluster failures at scale and define new attribution capabilities to close them, including how an NCCL timeout gets attributed to hardware, system software, or application, and which signals belong in a public API vs. an entitled runtime extension. • Work directly with cloud partners and cluster operators to translate their operational triage requirements into platform capabilities. • Drive community engagement strategy for the open source attribution project, including contribution models, ecosystem partnerships, and developer adoption. What We Need To See: • 12+ years total experience. • 5+ years in product management, solutions architecture, software engineering, or site reliability engineering on technical infrastructure products. Technical depth on GPU or AI infra is required. • Familiarity with GPU failure modes: XID codes, ECC correctable and uncorrectable errors, DCGM metrics, NVLink and PCIe health signals. • Technical understanding of why AI workloads fail: GPU hardware failure modes, distributed training failure patterns (NCCL timeouts, framework errors, hardware vs. system software vs. application fault distinction), and the telemetry systems (DCGM, sysfs, dmesg) that surface these signals. • Demonstrated ability to work with senior technical customers and translate operational requirements into product decisions. • Strong written and verbal communication across technical and non-technical audiences. • Bachelor's degree in Computer Science or equivalent experience. Ways To Stand Out From The Crowd: • Former SRE or infra engineer on GPU clusters. You have triaged job failures, parsed XID logs, and understand the cross-team deep dives that happen when a multi-thousand-GPU training run fails. • Deep familiarity with GPU scheduling, topology-aware placement, or multi-tenant GPU cluster management: Slurm prolog/epilog, Kueue, or Kubernetes-native scheduler environments. • Experience building products that produce structured diagnostics or attribution output: health verdict schemas, log analysis pipelines, rules-engine systems, or observability APIs at data center scale. • Practical experience delivering or contributing to an open-source infrastructure project, including managing contributor engagement on GitHub. NVIDIA is widely considered one of the technology world’s most d