Senior AI/ML Engineer - Model Serving and Optimization
Design and implement high-throughput, low-latency model serving solutions using PyTorch and TensorFlow. Optimize model performance on NVIDIA GPU hardware. Develop efficient Docker images for model and server deployment.
Key Highlights
Technical Skills Required
Job Description
Required Skillsets
AI/ML Domain Expertise
- Deep understanding of the AI/ML domain, with the core effort centered around model performance and serving, rather than general infrastructure.
- Expertise in PyTorch and TensorFlow: Proven ability to work with and troubleshoot model-specific dependencies, logic, and graph structures within these major frameworks.
- Production Inference Experience: Expertise in designing and implementing high-throughput, low-latency model serving solutions.
- Specialized Inference Servers: Mandatory experience with high-performance inference servers, specifically including vLLM, or similar dedicated LLM serving frameworks.
- GPU Optimization: Demonstrated ability to optimize model serving parameters and infrastructure to maximize performance on NVIDIA or equivalent GPU hardware.
- Containerization (Docker): Proficiency in creating minimal, secure, and efficient Docker images for model and server deployment.
- Infrastructure Knowledge (Helpful, but Secondary): General knowledge of cloud platforms (AWS, Google Cloud Platform, Azure) and Kubernetes/orchestration is beneficial but the primary focus remains on model serving and optimization.
Similar Jobs
Explore other opportunities that match your interests