Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
FAR Labs is a unified platform for high-performance AI inference. It provides dedicated Inference and Model APIs with flexible workload placement across its infrastructure, a customer’s cloud or hardware, and shared GPUs.
задачи
Set the technical direction for the inference stack and benchmarking practices;
Improve the speed and efficiency of model serving;
Validate performance gains with credible benchmarks;
Lead the engineers building the platform;
Shape the technology, performance standards, and engineering culture behind FAR Labs.
требования
Production experience with LLM inference-serving systems;
Expertise with vLLM, SGLang, TensorRT-LLM, or similar stacks;
Strong knowledge of batching, paged attention, KV-cache management, and modern serving architectures;
GPU performance engineering experience with CUDA or Triton;
Solid understanding of GPU utilisation, memory bandwidth, and precisions such as FP8 and FP4;