Jobs / United States / Hewlett Packard Enterprise Company

Senior Inference SW Engineer

Hewlett Packard Enterprise Company · 🇺🇸 3 Locations

Sponsorship verdict

Sponsorship possible

One solid signal, not two — worth applying, and worth asking about sponsorship early.

  • Employer is on a government sponsor recordThe US Department of Labor certified 486 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 167 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
  • The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
  • No salary bar for this routeH-1B has no fixed salary bar: the employer must pay at least the prevailing wage for the role and area. Cap-subject employers enter a lottery weighted by wage level. Source: https://www.federalregister.gov/documents/2025/12/29/2025-23853/weighted-selection-process-for-registrants-and-petitioners-seeking-to-file-cap-subject-h-1b, rules effective 2026-02-27.
  • What Hewlett Packard Enterprise Company paid sponsored hires in similar roles59 certified filings for “Systems/Software Engineer III” (Software Developers) in CA: $149k–$206k, median $183k. Most were filed at wage level II (43%) — 2 lottery entries, ≈31% projected selection odds for cap-subject employers. Source: US Department of Labor LCA disclosure data (Oct 2025 – Jun 2026).
  • Confirmed live todayWhen a source last listed this job as open.

US H-1B: cap-subject employers enter a lottery weighted by wage level — Level I gets 1 entry, Level IV gets 4 (DHS projected selection odds ≈15% at Level I to ≈61% at Level IV). Universities and non-profit research employers are cap-exempt. The $100,000 fee for new petitions from abroad is currently blocked by a court order (appeal pending).

A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.

Start free →

Or apply yourself on the official page →

Sponsor Radar — Hewlett Packard Enterprise Company

486 H-1B filings certified since Oct 2025

The US Department of Labor certified 486 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 167 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).

Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Hewlett Packard Enterprise Company →

About the role

Senior Inference SW Engineer    This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office. Who We Are: Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE. Job Description: HPE's Private Cloud AI organization is seeking a Senior Software Engineer  to build and evolve the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The core engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will design and implement key components of that runtime – engine integration, batching, KV cache management, and distributed execution – together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US ; however, remote work options will be considered. Responsibilities ·       Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution ·       Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency ·       Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage ·       Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption ·       Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling ·       Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence ·       Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team Knowledge and Skills Required ·       Familiar with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals ·       Strong understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding ·       Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them ·       Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling ·       Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight ·       Familiar with debugging/

View the official posting →

Source: Employer career site (Workday) First seen: 2026-10-07 Last confirmed: 2026-10-07 How our data works → Report this job

Similar opportunities