Skip to content
Portrait of Yaiser Avila Rodríguez

Fribourg, Switzerland · B permit · Open to remote roles

Senior SRE who keeps large Kubernetes fleets
fast, cheap and reliable

Nearly 5 years running production across 300+ clusters on AWS, Azure, GCP and Linode. I turn manual operations into automated, self healing platforms.

CKA certified · 10× faster deployments · 27% lower AWS cost

I engineer toil out of existence, so your fleet ships faster, costs less to run and stays calm when it matters. 

What I deliver

Faster releases

GitOps with ArgoCD and health gated progressive rollouts made deployments 10× faster and saved 100+ hours of manual work.

Lower cloud bills

Cut AWS platform cost by 27% and EKS extended support cost by 82.5% with autoscaling, right sizing and a planned 1.24 to 1.29 upgrade.

A quieter on call

Automated OOM remediation removed 20+ hours of toil and 20 to 30 incidents every month.

Calm incident command

Owned EMEA on call and incident command through peak traffic events such as Black Friday and the Super Bowl.

How we would work together

  1. 01

    Intro call

    Thirty minutes on your stack, your pain points and the role. No slides, no pressure.

  2. 02

    Technical deep dive

    I walk your engineers through real case studies, with the numbers, the trade offs and what I would change.

  3. 03

    A plan for the first 90 days

    I bring a written plan covering reliability, cost and delivery, based on what I learned from you.

Selected work

Four projects, each with the problem, what I did and what changed.

10×

faster deployments

Hydrolix · 2024 to present

GitOps across the production fleet

Problem
Hundreds of clusters had no single, auditable delivery model, so drift and slow manual releases were the norm.
What I did
Led the ArgoCD adoption from scratch: label architecture for regional rollouts, control plane topology, and the operating guides the SRE team still uses.
Outcome
Every production cluster now ships through automated continuous delivery, and it became the base for all later automation.
  • ArgoCD
  • Kubernetes
  • Kustomize
  • Pulumi
  • GitHub Actions

Fleet scale

health gated rollouts

Hydrolix · 2024 to present

Progressive delivery platform

Problem
Fleet wide upgrades needed weeks of manual coordination with no safe way to batch, retry or roll back.
What I did
Co built a platform on Argo Workflows with regional targeting, automated pre and post health checks, live DAG view over WebSocket and scheduled rollouts.
Outcome
Health gated rollouts with one click rollback, adopted as the standard way to upgrade the fleet.
  • Go
  • React
  • Argo Workflows
  • ArgoCD
  • Prometheus

20 to 30

incidents a month eliminated

Hydrolix · 2024 to present

Automated OOM remediation

Problem
Out of memory incidents kept paging the on call engineer, month after month, with the same manual fix.
What I did
Designed a remediation system that reads the observed signals and surfaces the right GitOps change automatically.
Outcome
More than 20 hours of toil and 20 to 30 incidents a month removed from the rotation.
  • Kubernetes
  • Prometheus
  • ArgoCD
  • Go
  • Pulumi

27%

lower AWS platform cost

Triggle Spain · 2023 to 2024

AWS cost optimisation and EKS upgrade

Problem
Over provisioned resources and clusters stuck on a version in paid extended support, billed at six times the standard rate.
What I did
Led a team of 3 through autoscaling policies, cleanup jobs, right sizing, and a 1.24 to 1.29 upgrade across 4 clusters with Terraform and Velero backups.
Outcome
27% lower platform cost and 82.5% lower extended support cost, with better scalability after the upgrade.
  • AWS
  • EKS
  • Terraform
  • Velero
  • CAST AI
Earlier work4 more projects
  • CI/CD automation and 5× faster buildsTriggle Spain · 2023
  • Zero cost on call platform with Grafana OnCallTriggle Spain · 2024
  • Docker layer cache in S3, builds from 20 to 5 minutesatSistemas · 2023
  • Branch Inspector, 80% shorter CI buildsAccenture · 2022

What colleagues say

Recommendations from people I worked with, as written on LinkedIn.

  • I had the pleasure of managing Yaiser for around 18 months as part of our SRE team. During that time, he helped maintain the reliability of our multi-cloud Kubernetes estate, which grew to several hundred clusters across four cloud providers. He contributed significantly to our move towards GitOps and ArgoCD... Above all, Yaiser was always reliable, willing to go the extra mile when needed, and genuinely a great guy to work with.
    Photo of Jason Turner

    Jason Turner

    Technology leader, Hydrolix

    Managed Yaiser directly · September 2026

  • He has a strong understanding of site reliability engineering and takes a thoughtful, practical approach to keeping the systems running, improving automation and solving issues. Yaiser is someone you can count on especially when things get challenging.
    Photo of Kelvin Wee

    Kelvin Wee

    Global Services Director, Hydrolix

    Senior to Yaiser · October 2026

  • Yaiser stays calm under pressure, which makes him someone people naturally lean on when things get difficult. He's genuinely curious and digs into problems until he truly understands them, so his decisions are thoughtful and well reasoned. He communicates clearly and without ego.
    Photo of Tameem Akhtar

    Tameem Akhtar

    Site Reliability, Hydrolix

    Worked with Yaiser on the same team · October 2026

  • He's one of the best SREs I've come across. He's the guy you want around when things break: calm, sharp, and always focused on fixing the problem. He also made our systems quieter and more reliable, and explained things in a way everyone could follow.
    Photo of Guillermo Vignau

    Guillermo Vignau

    Solutions Architect, Hydrolix

    Worked with Yaiser on the same team · October 2026

  • What makes Yaiser truly valuable is his consistency. Regardless of the environment or the circumstances, he brings the same reliable, team first energy. Any SRE or infrastructure team would benefit greatly from having him on board.
    Photo of Nelson Jaime Galván

    Nelson Jaime Galván

    DevOps and SRE, Hydrolix

    Worked with Yaiser at Accenture and Hydrolix · May 2026

  • Yaiser has strong hands on experience with Kubernetes and distributed systems... One of his key strengths is observability. During incidents, Yaiser stays calm and focused, quickly works through complex problems, identifies root causes, and drives issues to resolution while maintaining clear communication with the team.
    Photo of Jose Augusto Schiavoni

    Jose Augusto Schiavoni

    DevOps Engineer, Platform Engineering, Hydrolix

    Worked with Yaiser on the same team · October 2026

  • Yaiser demonstrated impressive technical expertise, effectively maintaining both on premises and AWS cloud services. His ability to automate deployments and ensure system scalability was crucial to the success of our projects. He was also a great support to junior profiles.
    Photo of Alex Curiman

    Alex Curiman

    Tech Lead, Epam Neoris

    Managed Yaiser directly at atSistemas · August 2024

  • He is one of the most skilled DevOps Engineers I have ever worked with. Yaiser excels not only in technical expertise but also in leadership and collaboration. He is also a fast learner, and found a way to educate himself with new technologies that were needed to specific projects.
    Photo of Jose Antonio Alvarez González

    Jose Antonio Alvarez González

    Java Software Developer, Triggle

    Worked with Yaiser on the same team · September 2024

Side projects and my own infra

What I build and run for myself, with the same habits I use at work: monitoring from outside, least privilege and infrastructure I can rebuild.

Running daily

Discord to Claude Code bridge

Drive a coding agent from a Discord thread

A TypeScript bot that opens a thread per task and streams Claude Code's work back while it runs on my machine. It runs as a launchd agent under my own user, limits its reach to one workspace with a Seatbelt sandbox, and replaced an earlier design that ran as root and buffered output.

  • TypeScript
  • launchd
  • Seatbelt
  • Redis
  • Cloudflare Tunnel

Running daily

Heartbeat monitoring, from outside the machine

A watcher that can notice when the machine dies

A host cannot report its own death, so the alarm lives elsewhere. A reporter on the Mac edits one pinned Discord message every 5 minutes and posts a beat to a Cloudflare Worker. A cron in the Worker checks the last beat in KV and alerts through a webhook, then posts a recovery line.

  • Cloudflare Workers
  • Workers KV
  • Cron triggers
  • Discord webhooks
  • launchd

Built in 2024

LinkedIn post scheduler

Personal brand automation on the edge

A small CRUD interface for post drafts and a Cloudflare Worker that reads them from a Turso SQLite database and publishes through the LinkedIn API at planned times.

  • React
  • Node.js
  • Turso
  • Cloudflare Workers
  • Wrangler

Private demo on request

Écoute, a French listening lab on local AI

Learn French by listening, with no data leaving my Mac

You paste a French text and a neural voice (Kokoro) reads it sentence by sentence. A local model (Ministral 3 8B on Ollama) prepares the lesson: translation, vocabulary in context, comprehension questions and a chat tutor. In dictation mode the diff is computed in code and the model only explains the differences. It runs on an Apple M3 Pro with 18 GB of memory, as a small Node server with no npm dependencies, published through a Cloudflare Tunnel behind Cloudflare Access.

  • Ollama
  • Kokoro TTS
  • Node.js
  • Cloudflare Tunnel
  • Cloudflare Access
Request demo access

Experience

  1. Site Reliability Engineer at Hydrolix

    Jul 2024 to present

    Remote, Switzerland · via Oyster HR

    Petabyte scale streaming analytics platform. Multi cloud Kubernetes fleet of 300+ clusters and 20,000+ pods.

    • Introduced the GitOps layer on ArgoCD and co built the progressive delivery platform behind 10× faster deployments.
    • Moved onboarding from VMs and manual steps to automated, secure customer onboarding with Argo Workflows, External Secrets and AWS Secrets Manager.
    • Built the global fleet monitoring, extended Prometheus to 30+ unmonitored environments, and owned EMEA on call through Black Friday and the Super Bowl.
  2. Senior DevOps Engineer at Triggle Spain

    Jul 2023 to Jul 2024

    Spain

    Tourism PaaS on AWS and Azure. Led a DevOps team of 3.

    • Cut AWS platform cost by 27% and accelerated deployments 5×.
    • Upgraded 4 EKS clusters from 1.24 to 1.29, saving 82.5% of extended support cost.
  3. Senior DevOps Engineer at atSistemas Consulting

    Feb 2023 to Jul 2023

    Spain

    Airline customer, on premise and AWS services.

    • Automated deployments and mentored 2 junior SREs.
  4. Junior SRE / DevOps at Accenture

    Feb 2022 to Feb 2023

    Spain

    Banking customer.

    • Cut CI build times by 80% with a GitLab API Branch Inspector tool in Python.

Stack and certifications

Daily tools

Orchestration and delivery
KubernetesArgoCDArgo WorkflowsHelmKustomize
Infrastructure as code
PulumiTerraformCloudFormationAnsible
Observability
PrometheusGrafanaOpenTelemetryVector
Secrets
External Secrets OperatorAWS Secrets Manager
Languages
GoPythonBashTypeScript
Clouds
AWSAzureGCPLinode

Certifications

  • Certified Kubernetes Administrator

    The Linux Foundation, 2026

    Certified
  • AWS Certified Cloud Practitioner

    Amazon Web Services

    Certified
  • Claude Code 101 and AI Fluency

    Anthropic Academy, 2026

    Training

Education: Master in Software Development (Assembler Institute), MSc Plant Breeding, MSc Plant Biology, BSc Biological Sciences.

Questions recruiters ask

Where are you based and can you work in Switzerland?

I live in Fribourg and hold a valid B permit, so no sponsorship is needed.

Do you work remotely?

Yes. I have worked fully remote on a global team since 2024 and I am on Central European time.

Are you looking for full time roles only?

Full time first. I will consider contracts when the scope is clear and the work is substantial.

What stack do you know best?

Kubernetes with ArgoCD and GitOps, Pulumi and Terraform, Go for tooling, and Prometheus with Grafana for observability.

How big are the environments you have run?

More than 300 production clusters and 20,000 pods across AWS, Azure, GCP and Linode.

Which certifications do you hold?

Certified Kubernetes Administrator (2026) and AWS Certified Cloud Practitioner, plus Anthropic Academy training on AI tooling.

Which languages do you speak?

Spanish is my native language, my English is fluent (C1) and used daily at work, and my French is A2 to B1 with courses under way, aiming for B1.

When can you start?

Notice period to be confirmed. Book a call and I will give you an exact date.

Let us talk about your fleet

Thirty minutes, no preparation needed. I am based in Switzerland, hold a B permit and work remotely.

Prefer email? [email protected]