Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
The role focuses on maintaining and improving the reliability, scalability, and security of infrastructure and CI/CD environments, including L3 production support, Kubernetes, automation, monitoring, incident management, and secure software delivery.
задачи
Provide L3 technical support for complex infrastructure, application runtime, and deployment issues;
Conduct root cause analysis and implement preventive measures;
Design, maintain, and continuously improve CI/CD pipelines using tools such as GitLab CI;
Manage and troubleshoot Kubernetes and other distributed environments;
Automate infrastructure provisioning, configuration management, and routine operational tasks using Infrastructure as Code and scripting;
Develop and improve monitoring, logging, tracing, and alerting;
Establish and maintain secure practices for secrets management, access control, configuration management, and service discovery;
Improve deployment, rollback, and recovery procedures;
Integrate automated security checks and quality controls into the software delivery lifecycle with security and development teams;
Collaborate with development and platform teams to assess production readiness;
Validate system resilience and disaster recovery procedures;
Maintain technical documentation and runbooks;
Share knowledge and mentor engineers to strengthen the team’s operational capabilities.
требования
Strong hands-on experience in DevOps engineering and production support;
Advanced knowledge of Linux administration, networking, and troubleshooting;
Practical experience with Kubernetes, Docker, GitLab CI, and Argo CD;
Proficiency in Bash or Python scripting;
Understanding of GitOps, secure software delivery, and secrets management;
Knowledge of high availability, disaster recovery, capacity planning, and DevOps practices;
Strong analytical thinking and structured problem-solving skills;
A high sense of ownership and accountability;
Ability to remain calm and make sound decisions during incidents;
Clear communication and effective collaboration across teams;
A proactive approach to identifying risks and improving processes;
Strong prioritization skills and attention to detail.