Walid El Sayed Aly

Freelance Cloud & Platform Engineer for AI Workloads. I build production-ready, GDPR-compliant GenAI platforms on Kubernetes — vLLM/LiteLLM for LLM serving, NVIDIA GPU Operator and Karpenter for GPU scheduling, Qdrant for RAG, EU AI Act-aware architectures for Java/Spring enterprises. Based in Limburg an der Lahn, serving Frankfurt am Main, Mainz, Koblenz, Hessen — and remote across Germany.

Focus

Platform Engineering for AI Workloads

Most companies I speak to have already run their GenAI pilot. It worked in a notebook, everyone was impressed, and then somebody asked what happens when it has to serve five hundred employees, keep customer data inside the EU, and survive an audit. That is usually where the project stops.

That gap is what I work on. Not model training and not prompt engineering — the platform underneath: inference that stays up, GPUs that cost less than the team using them, retrieval that runs on infrastructure you control, and enough logging that a compliance officer will sign it off.

I am a freelance Cloud & Platform Engineer based in Limburg an der Lahn, working with companies across Frankfurt am Main, Wiesbaden, Mainz and Koblenz — and remote throughout Germany. Fifteen years of enterprise Java and cloud-native infrastructure came before the AI part, which is probably why I think first about what operations will have to live with rather than what demonstrates well.

  • 01

    LLM Serving on Kubernetes

    Self-hosted inference with vLLM and LiteLLM, packaged as Helm charts. Multi-model routing, token quotas and API-key management — the parts that turn a running model into something you can hand to a development team.

  • 02

    GPU Scheduling & Cost

    NVIDIA GPU Operator, MIG partitioning and Karpenter on EKS or GKE. Spot strategies and honest right-sizing, so idle accelerators stop quietly draining the budget between demos.

  • 03

    RAG Infrastructure

    Qdrant, embedding pipelines and reranker models running inside your own network. Air-gapped setups included — that requirement comes up far more often in German industry than most vendor demos assume.

  • 04

    Observability & FinOps

    Prometheus and Grafana for GPU and token metrics, Langfuse for tracing. Cost per tenant, per model and per request becomes something you can answer instead of estimate.

  • 05

    EU AI Act & GDPR

    Risk classification, documented data flows and audit logging for self-hosted models. Written so the legal side can read it and the engineering side can implement it.

  • 06

    Spring AI Integration

    Connecting GenAI to the Java estate you already run. Spring Boot, Spring AI and custom gateways — extending the existing system rather than proposing a rewrite.

Approach

How I work

  • 01

    Boring infrastructure wins

    I would rather hand over a plain Helm chart that survives a year of production than a clever setup that breaks the first time someone adds a second GPU node.

  • 02

    Written down, in public

    Terraform modules, Helm charts and write-ups end up on this site and on GitHub — including the attempts that did not work. It is easier to judge someone by their notes than by their slides.

  • 03

    Compliance belongs in the design

    In insurance, banking and public-sector work, "we will sort out the audit later" is how projects quietly die. Data flows and logging go into the architecture from the start.

  • 04

    Handover, not dependency

    Your team should be able to operate the platform once I am gone. If they still need me for routine changes, the job is not finished.

Background

Where this comes from

Diplom-Informatiker, University of Applied Sciences Worms. Since 2013: Software Developer at 1&1 Internet in Montabaur, Lead Software Developer at Trusted Shops in Cologne, Lead Developer and Product Owner at Deutsche Bahn / DB Systel in Frankfurt, and since 2022 Software Architect at LR Health & Beauty.

Along the way: roughly twenty cloud migration and DevOps transformation projects, Kubernetes and Helm in production, GitLab CI and GitHub Actions pipelines, Terraform on AWS, and a great deal of Spring Boot. On several of those, deployment times came down by up to thirty percent.

The AI specialisation is newer, and I would rather say so than pretend otherwise. I am building it in the open — publishing what I learn about vLLM, GPU operators, RAG and the EU AI Act as I go, here and on GitHub. That way you can read the work instead of taking my word for it.

Work

Things I build in public

CICDX — a SaaS platform that converts CI/CD pipelines between Jenkins, GitHub Actions, GitLab CI, CircleCI and Azure Pipelines automatically. Migrations that normally eat weeks of an engineer's time, done as a translation problem. Currently in early access.

Alongside it: Helm charts and Terraform modules for the AI platform stack, plus write-ups of what happened when I ran them — GPU nodes that would not schedule, embeddings that were slower than the model, cost surprises worth warning people about.

Contact

Let's talk

If you are weighing up self-hosted LLM serving, trying to get GPU spend under control, or need a GenAI setup that will hold up to a GDPR or EU AI Act review, send a short description of where you are. No specification needed. I will tell you honestly whether it is something I can help with — and if it is not, I will say so.

hi@wesa.dev · Limburg an der Lahn · on-site across Frankfurt am Main, Wiesbaden, Mainz, Koblenz and Hessen · remote throughout Germany · German and English.