Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
The company operates an iGaming platform and focuses on reliable, scalable, observable and production-ready technology services.
задачи
Design, build and operate reliable, scalable infrastructure for the iGaming platform on AWS using Infrastructure as Code;
Own platform reliability across infrastructure and application services, including Kubernetes, networking, service health and performance;
Define and maintain SLIs and SLOs and establish reliability targets for critical services;
Build and improve end-to-end observability across metrics, logs and distributed tracing, including dashboards and actionable alerting;
Investigate and resolve production incidents end-to-end, perform root cause analysis and implement permanent reliability improvements;
Design, build and maintain CI/CD and GitOps workflows with automated validation, safe deployment strategies and reliable rollback mechanisms;
Work directly with application code when required by adding instrumentation, improving health checks, troubleshooting services and collaborating with developers;
Own and improve infrastructure for networking, load balancing, CDN, DNS, TLS and traffic management;
Identify opportunities to automate repetitive operational tasks and reduce manual intervention;
Improve security, performance, scalability, efficiency and cost predictability across the platform;
Collaborate with software engineers on architecture, scalability, resilience, failure scenarios and production readiness;
Contribute to technical architecture and identify potential reliability issues before production;
Develop and maintain runbooks, operational procedures and technical documentation;
Build and maintain a technical backlog of infrastructure and reliability improvements and drive initiatives through to completion;
Use modern AI tools and AI agents in daily engineering workflows for investigation, debugging, automation, code development and problem-solving.
требования
4+ Years of experience in SRE, DevOps, Platform Engineering or Infrastructure Engineering;
2+ Years of hands-on experience running Kubernetes in production;
Strong hands-on experience with AWS and cloud infrastructure, including compute, networking, IAM, storage, databases, load balancing, CDN and managed services;
Strong experience with Kubernetes, Terraform and Helm;
Practical experience with Prometheus, Grafana and OpenTelemetry or equivalent technologies;
Strong Linux and networking knowledge, including DNS, TLS, routing and container runtimes;
Experience building and maintaining CI/CD pipelines and GitOps workflows;
Ability to read, understand and modify application code when required;
Strong scripting and automation skills, particularly with Bash and/or Go;
Good understanding of distributed systems, scalability, high availability and common failure modes;
Strong troubleshooting and analytical skills across multiple layers of a technology stack;
Experience with modern AI-powered engineering tools and agents and willingness to use them in daily workflows;
Strong ownership and continuous-improvement mindset with the ability to build a technical backlog and drive initiatives forward;
Comfortable collaborating with software engineers and contributing to technical and architectural discussions;
Fluent Ukrainian or Russian is a must;
English B1 level or higher is a must;
Nice to have: Experience with GitHub Actions, ArgoCD or Flux.
условия
Hybrid freedom with an office and remote-work mix;
Flexible working hours;
21 Vacation days and 7 sick days;
Birthday half-day off;
Snacks and drinks provided;
Brand New MacBook provided;
Real learning, mentorship and fast skill development;