Jobs / United States / Nvidia Corporation
Senior DevOps Engineer, AIOps
Nvidia Corporation · 🇺🇸 2 Locations
Sponsorship verdict
Sponsorship possible
One solid signal, not two — worth applying, and worth asking about sponsorship early.
- Employer is on a government sponsor recordThe US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
- The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
- No salary bar for this routeH-1B has no fixed salary bar: the employer must pay at least the prevailing wage for the role and area. Cap-subject employers enter a lottery weighted by wage level. Source: https://www.federalregister.gov/documents/2025/12/29/2025-23853/weighted-selection-process-for-registrants-and-petitioners-seeking-to-file-cap-subject-h-1b, rules effective 2026-02-27.
- What Nvidia Corporation paid sponsored hires in similar roles701 certified filings for “Engineer Senior Systems Software” (Software Developers) in CA: $173k–$214k, median $190k. Most were filed at wage level IV (82%) — 4 lottery entries, ≈61% projected selection odds for cap-subject employers. Source: US Department of Labor LCA disclosure data (Oct 2025 – Jun 2026).
- Confirmed live todayWhen a source last listed this job as open.
US H-1B: cap-subject employers enter a lottery weighted by wage level — Level I gets 1 entry, Level IV gets 4 (DHS projected selection odds ≈15% at Level I to ≈61% at Level IV). Universities and non-profit research employers are cap-exempt. The $100,000 fee for new petitions from abroad is currently blocked by a court order (appeal pending).
A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.
Or apply yourself on the official page →
Sponsor Radar — Nvidia Corporation
The US Department of Labor certified 2,374 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 394 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Nvidia Corporation →
About the role
NVIDIA is powering the world’s most advanced AI factories, where resilient infrastructure is essential to keep accelerated computing environments running at scale. The Agentic AIOps team is building a mission-critical observability and prediction platform - delivered as both a high-scale SaaS solution and a robust on-premises deployment for NVIDIA’s largest enterprise customers. As a Senior DevOps Engineer, you’ll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIA’s AI infrastructure. What You'll Be Doing: • Own the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations. • Build and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints. • Engineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory. • Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity. • Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations. • Strengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows. • Deliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening. • Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes. What We Need to See: • Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent experience. • 5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices. • Strong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting. • Strong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible. • Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices. • Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting. • Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting. • Sound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment. Ways To Stand Out From the Crowd: • Experience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost. • Familiarity with Temporal,