Jobs / Germany / Tensordyne

Forward Deployed Inference Engineer

Tensordyne · 🌍 Munich, Bavaria, Germany

Sponsorship verdict

No sponsorship evidence yet

No government record and no wording either way. Not a refusal — ask the recruiter.

  • No government sponsor record hereThis employer posted directly and does not match a government sponsor register.
  • The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
  • Can’t check pay against the visa rulesNo salary stated. Germany EU Blue Card needs at least €50,700 a year (shortage-occupation / recent-graduate: €45,934). Source: https://www.make-it-in-germany.com/en/visa-residence/types/eu-blue-card, rules effective 2026-01-01.
  • Confirmed live todayWhen a source last listed this job as open.

A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.

Start free →

Or apply yourself on the official page →

Why not apply?

Sponsorship unknown

No register record and no sponsorship wording in the posting. Worth asking the employer before investing significant time.

SponsorApply flags time-wasters so your applications go where they can land. These come from the posting's own wording — read the original listing to confirm. See better-fit alternatives →

Sponsor Radar — Tensordyne

Employer-posted — not on a register

This employer posted directly and does not match a government sponsor register.

Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Tensordyne →

About the role

About Tensordyne Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world's most demanding generative AI workloads. Our platform combines purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture into a tightly integrated system purpose built for large-scale AI inference. We work with hyperscalers, Neoclouds, frontier model developers, enterprises, and infrastructure partners operating at the leading edge of AI. As Tensordyne moves from system development into silicon bring-up, customer validation, beta deployments, and production rollout, we are building the technical customer organization that will sit directly between our engineering teams and the companies deploying the platform. Role summary We are looking for a Forward Deployed Inference Engineer who combines deep AI systems expertise with strong customer instincts. This person will own the path from a customer workload or model request to a technical result and, where needed, to an optimized model running successfully on Tensordyne hardware and software. The role sits at the intersection of model architecture, inference performance, systems optimization, developer tooling, and customer deployment. You will work hands-on with engineering while also acting as a technical bridge to Product, BizDev, Sales, and customers. What you will do • Turn customer workloads into fast, credible performance answers through profiling & benchmarking . Define relevant KPIs, compare against competitive baselines, and keep our evaluation methodology current with external benchmarks. • Model enablement & optimization: convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify performance bottlenecks, and work with compiler, runtime, kernel, and system teams to improve results. • Deployment / forward engineering: work directly with customers and partners on technical PoCs, integration, deployment, and debugging; translate requirements into measurable acceptance criteria for quality, latency, throughput, and other relevant KPIs. • Track profiling-to-hardware accuracy by continuously comparing profiling/simulation results with actual hardware deployments, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, KV-cache management, or adjacent systems to the owning teams. • Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements. Core qualifications • Strong hands-on experience with AI models and inference systems , especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models. • Strong Python and PyTorch skills and the ability to understand and modify model code. • Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization. • Strong problem-solving and communication skills, with the ability to drive ambiguous technical problems across team boundaries. • Proficiency in using AI-powered developer tools (e.g., Claude Code, Cursor). Strong pluses • Experience with LLM serving and deployment stacks such as vLLM, SGLang, or similar systems. • Experience working directly with customers or external technical partners . • Experience bringing models up on new accelerators or non-standard hardware , including performance debugging across framework/runtime/hardware boundaries. • Practical experience with production inference techniques or environments such as quantization, distributed inference, or Kubernetes . • Experience navigating and contributing to Rust codebases. Tensordyne Values • Think big. Pursue ambitious technical a

View the official posting →

Source: Arbeitnow feed First seen: 2026-10-06 Last confirmed: 2026-10-06 How our data works → Report this job

Similar opportunities