Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
Claven AI is building a Sovereign AI Infrastructure & Orchestration Platform that turns bare-metal or managed GPU clusters into cloud-native AI clouds. The platform supports inference, model fine-tuning, and GPU workload management at scale, with a focus on data sovereignty for enterprise and government customers in regulated industries.
задачи
Build and evolve core Go services that orchestrate AI/ML workloads on Kubernetes, including inference traffic routing, workload provisioning, reconciliation, and lifecycle management;
Write controllers, operators, admission webhooks, and custom resources using controller-runtime / kubebuilder, including GPU-aware scheduling and placement;
Integrate and tune the model-serving stack for cost, latency, and throughput, including multi-GPU tensor-parallel serving and model artifact lifecycle management;
Implement multi-tenancy, network policy, and workload isolation for data-sovereignty and air-gapped operating modes, including per-tenant isolation, quotas, and RBAC;
Own the platform's API contracts and partner with frontend engineers on backend/UI boundaries, including gateway and identity layers;
Help diagnose and resolve production issues across the stack during pilot customer onboarding;
Own backend services and Kubernetes machinery that turn GPU clusters into a multi-tenant AI platform;
Actively use Claude Code, Cursor, or similar agentic AI tools in the daily workflow;
Investigate unfamiliar systems with AI assistance and return with answers, prototypes, and opinions;
Take ownership of problems within the assigned area;
Maintain curiosity and learning velocity.
требования
Strong production-grade Go experience building and operating real services;
Deep Kubernetes knowledge, including the API and controller/reconciliation model;
Experience operating Kubernetes clusters and debugging incidents;
Strong Linux, networking, and systems fundamentals;
Hands-on experience with infrastructure-as-code, observability tooling, and CI/CD pipelines;
Ability to own API contracts and communicate clearly in writing;
Strong written English and clear, low-ego communication;
4+ Years of experience in backend, platform, infrastructure, or SRE roles is a useful soft floor, not a hard gate;
Nice to have: Kubernetes controllers, operators, admission webhooks, or custom resources using controller-runtime / kubebuilder; NVIDIA stack including CUDA, MIG, NCCL, or DCGM; GPU partitioning and virtualization including MIG, time-slicing, or MPS; production LLM serving with vLLM, TGI, SGLang, or similar; multi-tenancy, network policy, or compliance work in regulated industries; Python; API-gateway or identity/OIDC work; ML platforms, developer infrastructure, or internal PaaS-style products; usage-based metering and quota systems.
условия
Fully remote, globally;
Working hours should overlap meaningfully with the team;
Ownership of a critical platform layer from day one;
Direct work with the founders and a small, AI-native engineering team;
Real customer pipeline in regulated markets, including pilots;
Compensation is discussed during the hiring process and is competitive for the candidate's market.