Staff / Principal Machine Learning Engineer, Serving - Switzerland

Inworld • , , switzerland, , , switzerland, Switzerland • Posted June 08, 2026

Location , , switzerland, , , switzerland
Job Type Full-time
Category Ingenieurwesen und Technik
Posted June 08, 2026

Who We're Looking For

A year ago, reliably working agentic systems and sub‑second multimodal inference at scale barely existed. Nobody has a decade of experience here. So we're not screening for a resume template — we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood.

Experience We Find Useful

  • Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT‑LLM.

  • Model Acceleration . Hands‑on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding.

  • High‑Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimised Python. You know how to profile code and squeeze every ounce of performance out of NVIDIA GPUs.

  • Distributed Systems & Scaling. Experi...

Interested in this role?

Click the button below to start your application.

Apply Now