About ZAP51

I work on the layers people notice only when they fail.

I’m a Principal Architect. My work moves from data centers, Linux, networks, and storage through cloud platforms, reliability, security, and the accelerated systems now serving AI workloads.

One architecture. Seven layers.

I do not see these as separate skill boxes. Each layer constrains the next, and good architecture respects the entire path from hardware to workload.

07
Workloads

Accelerated AI

Designing the compute and serving path for emerging AI workloads.

NVIDIA AIAcceleratorsvLLMLarge language modelsGPU systems
06
Application plane

Services and software

Building interfaces, automation, and integrations that make infrastructure usable.

REST APIsWeb servicesPythonJavaScriptGitOpen source
05
Platform plane

Cloud orchestration

Private and hybrid cloud platforms designed around workload, tenancy, and operational needs.

OpenStackApache CloudStackRed Hat OpenStackKubernetesOpenShiftAWSAlibaba Cloud
04
Control plane

Virtualization

Compute, software-defined networking, and cluster services beneath the cloud layer.

VMware vSphereNSXAVI Load BalancerTanzuProxmoxWindows Server
03
Data plane

Storage and networks

Distributed data, object services, routing, proxies, and security boundaries designed together.

CephIBM Spectrum ScaleObject storageNetworkingFortigateLoad balancingApplication security
02
Operating layer

Reliability and security

Keeping complex systems observable, recoverable, secure, and understandable under pressure.

Site reliability engineeringLinuxAnsibleSystem administrationCybersecurityIncident handlingEscalation
01
Foundation

Physical infrastructure

The power, placement, capacity, and operational realities every higher layer inherits.

Data centersComputeCapacityTechnical support

Architecture that survives contact with reality.

Technical depth matters, but so do judgment, communication, and staying present when a design becomes a production system.

01 · Architect

Start with constraints

Map failure domains, trust boundaries, capacity, dependencies, and operator experience before selecting technology.

02 · Operator

Stay for the incident

Use outages and escalations as design feedback, then automate the lesson instead of relying on memory.

03 · Researcher

Test the assumption

Read the source, reproduce the behavior, compare evidence, and separate a useful conclusion from a familiar explanation.

04 · Leader

Make the system legible

Build teams through clear decisions, shared context, useful documentation, and space for engineers to challenge assumptions.

AI needs infrastructure thinking.

Large model serving is not only a model problem. It is accelerator scheduling, memory, network movement, storage, observability, APIs, and reliability. That intersection is where my platform background becomes most useful.

NVIDIA AIvLLMAcceleratorsKubernetesObject storageNetworkingSRE
Explore the toolbox Read the blog GitHub Gitea SSH public key