Если вы раньше входили через Google, сбросьте пароль для своей Gmail-почты через кнопку «Забыли пароль?» на экране входа. Затем войдите по email и новому паролю.
Если аккаунта ещё нет, зарегистрируйтесь с Gmail-почтой, после подтверждения почты мы предложим задать пароль.
Что нового
Загружаю обновления...
Что нового
Загружаю обновления...
Работа найдется быстрее с подпискойКандидат найдётся быстрее с подпиской
Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
The role supports a large US hedge fund by keeping its critical Linux-based research and platform infrastructure reliable, secure, and efficient. The infrastructure includes Linux servers, Kubernetes platforms, scheduler-backed compute environments, and user access services.
задачи
Operate and support Linux servers and shared infrastructure used by research and platform teams
Troubleshoot production issues involving system performance, availability, access, configuration, and networking
Support Kubernetes and other container-based platforms, including node health, service behavior, and rollout activities
Support scheduler-backed compute environments such as Slurm, including node readiness, maintenance, and incident recovery
Improve user-facing Linux access services such as SSH, shared shell environments, and session-based platforms
Manage OS lifecycle work, including provisioning, patching, kernel and package updates, and hardening
Build scripts and automation in Python, Bash, or similar tools to reduce manual work and improve reliability
Use configuration management and version-controlled workflows to implement infrastructure changes safely
Enhance monitoring, alerting, documentation, and operational processes
Participate in incident response and occasional on-call support
требования
3+ Years of experience in Linux systems engineering, SRE, DevOps, or infrastructure support
Strong Linux administration skills, including systemd, package management, permissions, filesystems, log analysis, and performance troubleshooting
Good understanding of networking fundamentals, including DNS, NTP/PTP, routing, and general host connectivity
Experience with automation, scripting, and operational tooling
Familiarity with Kubernetes, virtualization, or clustered platforms
Experience with configuration management or infrastructure-as-code tools such as Ansible, Salt, or Terraform
Ability to troubleshoot production issues methodically and communicate clearly during incidents
Experience with Git-based workflows and maintainable documentation
Hands-on, practical problem solver with a strong ownership mindset
Comfortable working close to production and balancing support with continuous improvement
Collaborative communicator who works well across compute, storage, networking, and application teams
English: B2 Upper Intermediate
Будет плюсом:
Experience with Slurm, HPC-style environments, GPU infrastructure, or researcher-facing Linux platforms, working familiarity with shared storage clients such as NFS, autofs, or GPFS / IBM Storage Scale from a host and application perspective, experience with observability tools such as Prometheus, Grafana, or equivalent platforms, exposure to identity and access services such as LDAP, Kerberos, SSSD, or PAM, exposure to on-premises datacenter operations, hardware lifecycle support, or vendor escalations, interest in using AI/ML techniques for infrastructure optimization, anomaly detection, or predictive operations