O

Applied Research Engineer, LLM Compression

orient path_ Austria
Visa Sponsorship Relocation
Apply
AI Summary

Develop and ship new compression methods for Large Language Models, focusing on structural pruning and quantization. Own the technical scope from prototyping to production integration, optimizing models for cloud and edge deployment. Requires a PhD in CS/ML with published work on quantization/pruning and production-grade Python experience.

Key Highlights
Applied research role focused on shipping compression algorithms rather than writing papers.
Work across the stack including pruning, quantization, and evaluation infrastructure.
Hybrid work location in Vienna with visa sponsorship and relocation support.
Key Responsibilities
Improve and extend the structural pruning algorithm to new architectures such as MoE, multimodal, and vision-language models.
Combine pruning with quantization techniques (NVFP4/FP8/INT4, sub-4 bit mixed precision) in the compression pipeline.
Expand and improve the model retraining pipeline using SFT, GKD, DPO, and GRPO.
Compress customer models (Llama, Qwen, Gemma, and proprietary fine-tunes) for cloud and edge deployment.
Perform hardware-aware optimization for different accelerator targets including A100/H100/B300 and edge hardware.
Technical Skills Required
Python LLM Compression Quantization
Benefits & Perks
Visa sponsorship
Relocation support
Hybrid work in Vienna
Nice to Have
Open-source contributions to ML infrastructure (vLLM, llama.cpp, transformers, TensorRT-LLM, bitsandbytes, GPTQ/AWQ implementations).
Experience with MoE architectures or multimodal models (Qwen Omni).
Background in kernel optimization.

Job Description


At Orient Path, we partner with innovative deep-tech companies building the next generation of AI infrastructure.


About the Company

A deep-tech software startup of 6 people developing algorithms for LLM compression and optimization, founded in early 2025 by two former quantum physicists. The company raised a €3.5M Seed round in 2026, and plans to expand the team and compression offerings to Silicon Vendors, AI Enterprises, OEMs, and Cloud Providers.

We believe the next wave of AI adoption will be driven by compact, highly efficient models optimized for specific use cases, rather than large general-purpose cloud models.


About the Role

This is an applied research role. You'll develop new compression methods and ship them, not write papers about them. The cycle is short: read the literature, prototype, benchmark on real models, integrate into our pipeline, iterate with customers running compressed models in production.

You'll own significant technical scope from day one. Expect to work across the stack: pruning algorithms, quantization, evaluation infrastructure, and the production code that customers actually use.


Your impact

  • Improve and extend the structural pruning algorithm to new architectures (MoE, multimodal, vision-language).
  • Combine pruning with quantization (NVFP4/FP8/INT4, sub-4 bit mixed precision) in the compression pipeline.
  • Expand and improve the model retraining pipeline (SFT, GKD, DPO, GRPO).
  • Compress customer models (Llama, Qwen, Gemma, and proprietary fine-tunes) for cloud and edge deployment.
  • Hardware-aware optimization for different accelerator targets (A100/H100/B300 and edge hardware).


What we're looking for

  • PhD in computer science, machine learning, or equivalent.
  • Published work on quantization, pruning, or LLM training.
  • Production-grade Python code (not just Jupyter notebooks).
  • Experience taking a method from paper to a working system on real models.
  • Comfort working with LLMs, GPUs, and evaluating benchmarks.
  • You ship. You finish things.


Nice to have

  • Open-source contributions to ML infrastructure (vLLM, llama.cpp, transformers, TensorRT-LLM, bitsandbytes, GPTQ/AWQ implementations).
  • Experience with MoE architectures or multimodal models (Qwen Omni).
  • Background in kernel optimization.


Compensation & details

  • Hybrid in Vienna.
  • Visa sponsorship and relocation support available.
  • Full-time
  • Working language is English.



Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Silicon Austria Labs (SAL)

Austria

Senior IC-Design Flow Engineer

Programming
1w ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Silicon Austria Labs (SAL)

Austria
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

Nuki Home Solutions GmbH

Austria

Subscribe our newsletter

New Things Will Always Update Regularly