Jobs / Switzerland / Jobgether

AI Research Engineer (Kernel & Inference Optimization)

Jobgether · 🌍 Switzerland

Sponsorship verdict

No sponsorship evidence yet

No government record and no wording either way. Not a refusal — ask the recruiter.

  • No government sponsor record hereThis employer posted directly and does not match a government sponsor register.
  • The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
  • Can’t check pay against the visa rulesWe don’t have visa salary rules for this country yet.
  • Confirmed live todayWhen a source last listed this job as open.

A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.

Start free →

Or apply yourself on the official page →

Why not apply?

Sponsorship unknown

No register record and no sponsorship wording in the posting. Worth asking the employer before investing significant time.

SponsorApply flags time-wasters so your applications go where they can land. These come from the posting's own wording — read the original listing to confirm. See better-fit alternatives →

Sponsor Radar — Jobgether

Employer-posted — not on a register

This employer posted directly and does not match a government sponsor register.

Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Jobgether →

About the role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Switzerland. You will work at the intersection of AI research, systems engineering, and high-performance model inference. Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments. You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices. The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels. You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers. Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements. You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems. Accountabilities • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization. • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms. • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability. • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates. • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions. • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations. • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL). • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding. • Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads. • Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications. • Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies. • Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability. Requirements: • Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences. • Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch. • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices. • Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications. • Deep understanding of modern model-serving architectures, inference engines, and o

View the official posting →

Source: Arbeitnow feed First seen: 2026-10-01 Last confirmed: 2026-10-03 How our data works → Report this job

Similar opportunities