O

Senior Software Engineer (ML Infrastructure & Compression Pipeline)

ora computing • Austria
Visa Sponsorship Relocation
Apply
AI Summary

Own the design and optimization of ora computing’s core ML software stack, focusing on scalable compression pipelines, GPU performance, and robust library architecture. Refactor existing code into well-structured packages, build internal tooling, and set engineering standards for a growing team. Requires deep expertise in Python, GPU optimization, and production-grade software engineering.

Key Highlights
Design and refactor core ML libraries (pruning, quantization, retraining, evaluation) into scalable, well-scoped packages
Build internal tooling (CI, benchmarks, reproducible runs) to enable rapid, reliable development
Integrate compression pipelines with inference engines (vLLM, TensorRT-LLM, llama.cpp) and customer deployment targets
Set engineering standards for a growing team and hire against high technical bar
Key Responsibilities
Design and refactor core ML libraries (pruning, quantization, retraining, evaluation) into clean, modular packages
Build internal tooling (CI/CD pipelines, benchmarking frameworks, reproducible run environments) to support rapid development
Develop automated compression pipelines that account for target runtimes and optimize GPU performance
Integrate compression outputs with inference engines (vLLM, TensorRT-LLM, llama.cpp) and customer deployment targets
Establish engineering standards and best practices for a scaling team
Technical Skills Required
Python GPU Optimization (Memory Hierarchy, Kernels, Performance Bottlenecks) Software Architecture & Library Design
Benefits & Perks
€70–120k base salary + equity
Visa sponsorship and relocation support
Hybrid or fully remote work with English as the working language
Nice to Have
Open-source contributions to ML infrastructure (vLLM, llama.cpp, Transformers, TensorRT-LLM, PyTorch internals)
CUDA, Triton, or kernel-level optimization experience
Experience designing and shipping production-grade ML libraries used by other engineers
Familiarity with model serving and inference optimization

Job Description


THE ROLE

About The Role

You'll own how our software stack is built. Today the codebase reflects four people moving fast, it works, but it needs structure. Your job is to give it that structure: well-designed libraries, robust packages, environments, the kind of codebase that scales as we grow the team and ship more to customers.

This is not a glue-code role. You'll work between the algorithm and inference layer: designing a compression pipeline that is fully automated and takes target runtimes into account. You'll design the abstractions our compression pipeline runs on and make them fast.

What You'll Do

Your impact

  • Design and refactor our core libraries — pruning, quantization, retraining, evaluation — into clean, well-scoped packages.
  • Build the internal tooling that lets the team move quickly without breaking things — CI, benchmarks, reproducible runs.
  • Integrate our compression output with inference engines (vLLM, TensorRT-LLM, llama.cpp) and customer deployment targets.
  • Set the engineering bar for the team as we hire.

What You Bring

What we're looking for

  • Bachelor's/Master's in computer science or equivalent, plus 2+ years of professional software engineering.
  • Strong opinions about code design. You know what a well-structured library looks like and why.
  • GPU experience — memory hierarchy, kernels, what bottlenecks performance — even if you don't write CUDA daily.
  • Production-grade Python. You write code others can read, extend, and trust.
  • You finish things and you care about the codebase you leave behind.

NICE TO HAVE

  • Open-source contributions to ML infrastructure (vLLM, llama.cpp, transformers, TensorRT-LLM, PyTorch internals).
  • CUDA, Triton, or kernel-level work.
  • Experience designing a library from scratch that other engineers ended up using.
  • Familiarity with model serving and inference optimization.

PRACTICAL

  • Vienna-based. Hybrid or fully remote.
  • Working language is English.
  • We sponsor visas and support relocation.
  • Compensation: €70–120k base + equity. Austrian minimum disclosed per Kollektivvertrag: €45,738/year.
  • You'll set the engineering standards we hire against next.

Similar Jobs

Explore other opportunities that match your interests

Senior AI/GenAI Solutions Architect (Cloud & Distributed Systems)

Programming
•
1d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

INNIO Group

Austria

Head of Growth

Programming
•
4d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

clera

Austria
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

ora computing

Austria

Subscribe our newsletter

New Things Will Always Update Regularly