Jobs / United States / Nvidia Corporation

Senior Deep Learning Software Engineer, Inference and Model Optimization

Nvidia Corporation · 🇺🇸 2 Locations

Sponsorship verdict

Sponsorship possible

One solid signal, not two — worth applying, and worth asking about sponsorship early.

  • Employer is on a government sponsor recordThe US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
  • The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
  • No salary bar for this routeH-1B has no fixed salary bar: the employer must pay at least the prevailing wage for the role and area. Cap-subject employers enter a lottery weighted by wage level. Source: https://www.federalregister.gov/documents/2025/12/29/2025-23853/weighted-selection-process-for-registrants-and-petitioners-seeking-to-file-cap-subject-h-1b, rules effective 2026-02-27.
  • What Nvidia Corporation paid sponsored hires in similar roles701 certified filings for “Engineer Senior Systems Software” (Software Developers) in CA: $173k–$214k, median $190k. Most were filed at wage level IV (82%) — 4 lottery entries, ≈61% projected selection odds for cap-subject employers. Source: US Department of Labor LCA disclosure data (Oct 2025 – Jun 2026).
  • Confirmed live todayWhen a source last listed this job as open.

US H-1B: cap-subject employers enter a lottery weighted by wage level — Level I gets 1 entry, Level IV gets 4 (DHS projected selection odds ≈15% at Level I to ≈61% at Level IV). Universities and non-profit research employers are cap-exempt. The $100,000 fee for new petitions from abroad is currently blocked by a court order (appeal pending).

A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.

Start free →

Or apply yourself on the official page →

Sponsor Radar — Nvidia Corporation

2,374 H-1B filings certified since Oct 2025

The US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).

Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Nvidia Corporation →

About the role

NVIDIA is at the forefront of the generative AI revolution! The Algorithmic Model Optimization Team specifically focuses on optimizing generative AI models such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging from neural architecture search and pruning to sparsity, quantization, and automated deployment strategies. Our work includes conducting applied research to improve model efficiency as well as developing an innovative software platform (TRT Model Optimizer). Our software is used both internally across NVIDIA and externally by research and engineering teams alike developing best-in-class AI models. We are now looking for a Senior Deep Learning Software Engineer to develop and scale up our automated inference and deployment solution. As part of the team, you will be instrumental in pushing the limits of inference efficiency and large-scale, automated deployment. Your work will touch upon fundamental aspects of a typical machine learning stack including working in high-level frameworks like PyTorch and HuggingFace to developing and improving high-performance kernel implementations in CUDA, TRT-LLM, and Triton. This is an exceptional opportunity for passionate software engineers straddling the boundaries of research and engineering, with a strong background in both machine learning fundamentals and software architecture & engineering. What you’ll be doing: • Train, develop, and deploy state-of-the generative AI models like LLMs and diffusion models using NVIDIA's AI software stack. • Leverage and build upon the torch 2.0 ecosystem (TorchDynamo, torch.export, torch.compile, etc...) to analyze and extract standardized model graph representation from arbitrary torch models for our automated deployment solution. • Develop high-performance optimization techniques for inference, such as automated model sharding techniques (e.g. tensor parallelism, sequence parallelism), efficient attention kernels with kv-caching, and more. • Collaborate with teams across NVIDIA to use performant kernel implementations within our automated deployment solution. • Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities. • Continuously innovate on the inference performance to ensure NVIDIA's inference software solutions (TRT, TRT-LLM, TRT Model Optimizer) can maintain and increase its leadership in the market. • Play a pivotal role in architecting and designing a modular and scalable software platform to provide an excellent user experience with broad model support and optimization techniques to increase adoption. What we need to see: • Masters, PhD, or equivalent experience in Computer Science, AI, Applied Math, or related field. • 8+ years of relevant work or research experience in Deep Learning. • Excellent software design skills, including debugging, performance analysis, and test design. • Strong proficiency in Python, PyTorch, and related ML tools (e.g. HuggingFace). • Strong algorithms and programming fundamentals. • Good written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment. Ways to stand out from the crowd: • Contributions to PyTorch, JAX, or other Machine Learning Frameworks. • Knowledge of GPU architecture and compilation stack, and capability of understanding and debugging end-to-end performance. • Familiarity with NVIDIA's deep learning SDKs such as TensorRT. • Prior experience in writing high-performance GPU kernels for machine learning workloads in frameworks such as CUDA, CUTLASS, or Triton. Increasingly known as “the AI computing company” and widely considered to be one of the technology world’s most desirable employers, NVIDIA offe

View the official posting →

Source: Employer career site (Workday) First seen: 2026-10-10 Last confirmed: 2026-10-10 How our data works → Report this job

Similar opportunities