Applied Research Engineer - LLM Compression and Quantization
Develop and ship new compression methods for large language models. Own end-to-end pipeline from algorithm prototyping to production deployment. Requires PhD-level expertise in ML with production experience.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
THE ROLE
About The Role
This is an applied research role. You'll develop new compression methods and ship them, not write papers about them. The cycle is short: read the literature, prototype, benchmark on real models, integrate into our pipeline, iterate with customers running compressed models in production.
You'll own significant technical scope from day one. Expect to work across the stack: pruning algorithms, quantization, evaluation infrastructure, and the production code that customers actually use.
What You'll Do
Your impact
- —Improve and extend our structural pruning algorithm to new architectures (MoE, multimodal, vision-language).
- —Combine pruning with quantization (NVFP4/FP8/INT4, sub-4 bit mixed precision) in our compression pipeline.
- —Expand and improve our model retraining pipeline (SFT, GKD, DPO, GRPO).
- —Compress customer models (Llama, Qwen, Gemma, and proprietary fine-tunes) for cloud and edge deployment.
- —Hardware-aware optimization for different accelerator targets (A100/H100/B300 and edge hardware).
Looking to advance your Development & Programming career with relocation support? Explore Development & Programming Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
What we're looking for
- —PhD in computer science, machine learning, or equivalent.
- —Published work on quantization, pruning, or LLM training.
- —Production-grade Python code (not just Jupyter notebooks). You write code others can read and run.
- —Experience taking a method from paper to a working system on real models.
- —Comfort working with LLMs, GPUs, and evaluating benchmarks.
- —You ship. You finish things.
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- —Open-source contributions to ML infrastructure (vLLM, llama.cpp, transformers, TensorRT-LLM, bitsandbytes, GPTQ/AWQ implementations).
- —Experience with MoE architectures or multimodal models (Qwen Omni).
- —Background in kernel optimization.
- —Vienna-based. Hybrid or fully remote.
- —Working language is English.
- —We sponsor visas and support relocation.
- —Compensation: €70–120k base + equity. Austrian minimum disclosed per Kollektivvertrag: €43,456/year.
- —We don't require writing publications, but we support presenting work at venues when it fits the company and the project.
Similar Jobs
Explore other opportunities that match your interests