Artem Vaganovich

DevOps · SRE · Platform Engineering

Artem Vaganovich

Senior DevOps Engineer · SRE · Platform & Cloud Infrastructure

01

Professional summary

11+ years in production engineering: Kubernetes platforms, GitOps delivery pipelines, infrastructure as code, observability and on-prem/cloud reliability for highload systems — now with AI-assisted operations.

Senior engineer with 11+ years of hands-on production experience, focused on DevOps, SRE and platform engineering for high-load distributed systems. I build and operate the delivery and runtime platform: Kubernetes clusters, CI/CD pipelines, infrastructure as code, observability and incident response.

Strong background on both sides of the wall — I have designed and written the services I deploy, which makes capacity planning, performance tuning, database operations and production troubleshooting far more precise.

Extensive experience with Docker, Kubernetes, Helm, GitLab CI/CD, Terraform, Ansible, Linux, Nginx, Prometheus/Grafana, ELK/Loki and OpenTelemetry across cloud (AWS, Yandex Cloud) and on-premise environments with strict data-residency requirements.

Introduce AI into operations: agents and MCP tooling over internal infrastructure for incident triage, log analysis, runbook automation and pipeline generation — built on locally hosted and Russian LLM providers, with no foreign cloud dependencies.

Focused on reliability, zero-downtime delivery, cost efficiency and making the developer experience fast and safe.

Target roleSenior DevOps · SRE · Platform Engineer

02

Experience

Present

VSK

Senior Engineer · DevOps / SRE Practices

One of Russia's largest insurance companies

Own the delivery pipeline for mission-critical services: GitLab CI/CD with automated build, test, security scanning and staged rollout into Kubernetes, cutting release time and manual steps to a minimum.

Operate containerized workloads in Kubernetes — Helm charts, resource limits and autoscaling, health probes, zero-downtime rolling deployments, configuration and secrets management across environments.

Built observability: metrics and dashboards in Prometheus/Grafana, centralized logging, distributed tracing and alerting tied to real SLOs. Reduced mean time to detect and resolve production incidents.

Run performance and reliability work end to end: PostgreSQL tuning and connection pooling, Redis caching, queue throughput, capacity planning, load testing and post-incident hardening.

Introduced AI into operations — agents and MCP servers over internal systems for incident triage, log summarization, runbook and pipeline automation — built strictly on locally hosted and Russian LLM providers in line with security and data-residency policies.

KubernetesDockerHelmGitLab CI/CDLinuxNginxPrometheusGrafanaPostgreSQLRedisAI AgentsMCP
Previously

EPAM Systems

Senior Backend Engineer · Infrastructure & Cloud

Enterprise & US Marketplace Projects (including Walmart Labs)

Containerized and deployed large-scale marketplace and enterprise services, automating build and release pipelines and standardizing environments across teams.

Managed cloud deployments and infrastructure automation, including environment provisioning, configuration management, secrets handling and repeatable release procedures.

Operated data-layer infrastructure at scale: PostgreSQL, Redis, RabbitMQ and Elasticsearch clusters — indexing strategies, throughput tuning, and query and cluster performance optimization.

Handled production troubleshooting for high-traffic systems: profiling, log and metric analysis, resiliency patterns, caching and fault-tolerant communication between services.

Worked with distributed international teams on release planning, code review and continuous improvement of delivery and engineering practices.

DockerKubernetesCI/CDAWSLinuxPostgreSQLRedisRabbitMQElasticsearchNode.js
03

Selected projects

01

Kubernetes Delivery Platform

Unified delivery platform for dozens of services: from a git push to a verified production release without manual steps.

  • Templated GitLab CI/CD pipelines with build, tests, image scanning and staged promotion across dev/stage/prod.
  • Helm-based deployments with per-environment values, resource limits, autoscaling and health probes.
  • Zero-downtime rolling releases with automated rollback on failing health checks.
  • Standardized secrets and configuration handling across all services.
KubernetesHelmGitLab CI/CDGitOps
02

Observability & SRE Stack

End-to-end visibility for high-load production: metrics, logs, traces and alerting mapped to service-level objectives.

  • Prometheus metrics and Grafana dashboards per service and per business flow.
  • Centralized structured logging and distributed tracing for cross-service request analysis.
  • SLO-based alerting with actionable runbooks instead of alert noise.
  • Measurable reduction of mean time to detect and mean time to recover.
PrometheusGrafanaOpenTelemetryELK / Loki
03

Highload Infrastructure Operations

Keeping millions-of-users platforms fast and stable: capacity planning, database and cache operations, queue throughput and failure isolation.

  • PostgreSQL operations: query and index tuning, connection pooling, replication and backup/restore drills.
  • Redis caching layers and RabbitMQ queue tuning for peak traffic.
  • Elasticsearch cluster operations: sharding, indexing strategy and search latency control.
  • Load testing and capacity planning ahead of traffic peaks.
PostgreSQLRedisRabbitMQElasticsearch
04

AI-Assisted Operations

AI agents wired into internal infrastructure through Model Context Protocol to reduce toil in operations — running fully on-prem.

  • MCP servers exposing logs, metrics and deployment state as safe, read-only tools for agents.
  • Automated incident triage: correlation of alerts, logs and recent deployments into a first-pass diagnosis.
  • Generation and maintenance of pipelines, manifests and runbooks with AI assistance and human review.
  • Built on locally hosted and Russian LLM providers to satisfy data-residency and security policies.
AI AgentsMCPOn-Prem LLMAutomation
04

Core expertise

Key skills

Kubernetes · Docker · GitLab CI/CD · Terraform · Helm · Linux · Nginx · Prometheus / Grafana · PostgreSQL Ops · AI Ops Agents

Containers & Orchestration

  • Kubernetes (Deployments, HPA, Ingress)
  • Docker / Multi-stage builds
  • Helm charts
  • Service mesh basics
  • Autoscaling & resource tuning

Observability & SRE

  • Prometheus / Grafana
  • Loki / ELK logging
  • Distributed tracing (OpenTelemetry)
  • SLO / SLI & alerting
  • Incident response & postmortems

Reliability & Performance

  • Capacity planning
  • Load & stress testing
  • Zero-downtime deployments
  • Backup & disaster recovery
  • Production diagnostics

Linux, Network & Data

  • Linux administration
  • Nginx / reverse proxy / TLS
  • PostgreSQL & Redis operations
  • RabbitMQ / queues
  • Elasticsearch cluster ops

Security & Compliance

  • Least-privilege access & RBAC
  • Secrets & certificate rotation
  • Image scanning in pipelines
  • On-prem / data-residency setups

Infrastructure as Code

  • Terraform
  • Ansible
  • Docker Compose
  • Environment templating
  • Secrets management

CI/CD & GitOps

  • GitLab CI/CD pipelines
  • GitHub Actions
  • Blue-green & canary releases
  • Artifact & registry management
  • Automated quality gates

AI for Operations

  • AI agents for incident triage
  • Model Context Protocol (MCP) over infra tools
  • Log & alert summarization with LLM
  • On-prem / Russian LLM providers
  • Pipeline automation with AI

Cloud Platforms

  • Yandex Cloud
  • AWS
  • On-premise clusters
  • Hybrid infrastructure
05

Contact

Artem Vaganovich

Senior DevOps · SRE · Platform Engineer

Available for Senior DevOps / SRE / Platform Engineer roles