platform engineer for AI infrastructure
ориентир по рынку
вакансия
зп не указана
в среднем
328 556 ₽
мэтч
Загрузи резюме, чтобы видеть мэтчи с вакансией
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
The company provides AI infrastructure for large-scale model training and high-volume inference.
задачи
- Build and operate GPU-enabled Kubernetes infrastructure;
- Support distributed model training and large-scale inference workloads;
- Optimise GPU scheduling, allocation, utilisation, and performance;
- Build distributed compute environments using Ray;
- Develop platform tooling and automation in Python;
- Provision infrastructure using Terraform;
- Improve observability across GPU, Kubernetes, and workload layers;
- Troubleshoot performance bottlenecks across compute, memory, networking, and storage;
- Build self-service capabilities for ML and AI engineering teams;
- Improve platform reliability, scalability, and compute efficiency.
требования
- 4+ Years in ML Infrastructure, Platform Engineering, MLOps, HPC, or similar roles;
- Kubernetes;
- NVIDIA GPU infrastructure;
- CUDA ecosystem;
- Python;
- Terraform;
- Ray or comparable distributed compute technologies;
- Monitoring and observability;
- Strong understanding of Linux and distributed systems;
- Nice to have: NVIDIA GPU Operator, NCCL and distributed GPU communication, PyTorch distributed training, Triton Inference Server / vLLM, Prometheus / Grafana / OpenTelemetry, AWS, GCP, or Azure GPU infrastructure, Slurm or HPC environments, GPU capacity planning and cost optimisation, large-scale LLM training or inference.
условия
- Remote work across the European Union.
Если просят выйти из iCloud, прислать код из SMS, запустить или установить что-то, перевести деньги — не соглашайтесь: это мошенничество.
Про зарплаты
Анонимные данные по зарплатам и грейдам.
Можно сверить вилку с рынком.
Посмотреть зарплаты
статьи для DevOps-инженеров