Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
The company develops AI infrastructure for serving production AI models across GPU-enabled infrastructure, focusing on distributed systems, platform engineering, inference performance, GPU utilisation, autoscaling, latency, and reliability.
задачи
Build and operate Kubernetes infrastructure for production AI inference;
Deploy and optimise model-serving workloads using vLLM and NVIDIA Triton;
Improve GPU utilisation, throughput, and inference latency;
Design autoscaling strategies for dynamic AI workloads;
Build platform tooling and automation in Python;
Provision and manage infrastructure using Terraform;
Develop observability across models, GPUs, Kubernetes, and serving infrastructure;
Profile and troubleshoot inference performance bottlenecks;
Improve batching, concurrency, caching, and resource allocation strategies;
Build reliable deployment workflows for new models and model versions;
Partner with ML Engineers to move models efficiently into production.
требования
4+ Years in AI Infrastructure, ML Infrastructure, MLOps, Platform Engineering, or similar roles;
Strong understanding of Linux and distributed production systems;