Jobs / United States / Nvidia Corporation
Engineering Manager – AI Platform & SRE
Nvidia Corporation · 🇺🇸 US, CA, Santa Clara
Sponsorship verdict
Sponsorship possible
One solid signal, not two — worth applying, and worth asking about sponsorship early.
- Employer is on a government sponsor recordThe US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
- The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
- No salary bar for this routeH-1B has no fixed salary bar: the employer must pay at least the prevailing wage for the role and area. Cap-subject employers enter a lottery weighted by wage level. Source: https://www.federalregister.gov/documents/2025/12/29/2025-23853/weighted-selection-process-for-registrants-and-petitioners-seeking-to-file-cap-subject-h-1b, rules effective 2026-02-27.
- What Nvidia Corporation paid sponsored hires in similar roles113 certified filings for “Manager Systems Software” (Architectural and Engineering Managers) in CA: $214k–$248k, median $222k. Most were filed at wage level II (40%) — 2 lottery entries, ≈31% projected selection odds for cap-subject employers. Source: US Department of Labor LCA disclosure data (Oct 2025 – Jun 2026).
- Confirmed live todayWhen a source last listed this job as open.
US H-1B: cap-subject employers enter a lottery weighted by wage level — Level I gets 1 entry, Level IV gets 4 (DHS projected selection odds ≈15% at Level I to ≈61% at Level IV). Universities and non-profit research employers are cap-exempt. The $100,000 fee for new petitions from abroad is currently blocked by a court order (appeal pending).
A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.
Or apply yourself on the official page →
Sponsor Radar — Nvidia Corporation
The US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Nvidia Corporation →
About the role
Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused on designing, building, and operating large-scale production systems with exceptional reliability, efficiency, and availability. It combines software and systems engineering practices with expertise across distributed systems, networking, Kubernetes, public cloud, observability, capacity management, continuous delivery, and automation. As an Engineering Manager, you will lead a team of dedicated engineers responsible for building and operating resilient AI platform capabilities at enterprise scale. You will combine people leadership with strong technical judgment, helping the team translate ambiguous business and engineering challenges into a clear strategy and executable roadmap. You will partner across Cloud, Platform, Security, and AI/ML organizations to deliver reliable systems, improve developer productivity, and advance the use of AI agents and skills in platform operations. Our culture values diversity, intellectual curiosity, collaboration, and continuous learning. We encourage thoughtful risk-taking, blameless analysis, and shared ownership. You will create an environment in which engineers can do their best work, grow their careers, and make a meaningful impact. What you’ll be doing • Lead, develop, and grow a team of SRE, platform, and software engineers responsible for NVIDIA’s AI Platform Runtime and related production services. • Define the team’s technical strategy, priorities, and roadmap in alignment with broader product, platform, and business objectives. • Guide the design and delivery of highly available, scalable, secure, and resilient distributed systems that support enterprise AI agent products. • Drive the development of AI agents, AI skills, and intelligent automation for platform operations, incident response, troubleshooting, and remediation. • Establish measurable reliability goals and effective operational practices using service-level indicators, service-level objectives, error budgets, capacity models, operational health metrics, and production readiness reviews. • Improve engineering velocity and developer experience through self-service platforms, infrastructure-as-code, standardized delivery patterns, and automation. • Partner with product managers, architects, and leaders across Cloud, Security, Networking, Platform, and AI/ML teams to coordinate initiatives involving multiple functions. • Maintain a healthy balance among feature delivery, platform investment, operational work, reliability improvements, and technical debt reduction. • Lead the team through critical incidents and ensure that blameless postmortems result in clear ownership and durable corrective actions. • Recruit exceptional engineers and foster an inclusive, high-performing environment through coaching, feedback, career development, and thoughtful delegation. What we need to see • 10+ years of experience in Site Reliability Engineering, Platform Engineering, Software Engineering, Cloud Infrastructure, or a related technical field, including 3+ years managing or formally leading engineering teams responsible for complex production systems. • Technical foundation in distributed systems, Linux, networking, Kubernetes, and public cloud platforms such as AWS, Azure, or GCP. • Experience leading teams that build production software and automation using languages such as Python, Go, TypeScript, JavaScript, or Java. • Solid understanding of observability at scale, including OpenTelemetry, metrics, logs, distributed tracing, profiling, and operational analytics. • Experience applying SRE practices such as service-level objectives, error budgets, capacity and resource management, incident management, disaster recovery, and blameless postmortems.