Incident Recovery
Recovered Business-Critical MongoDB Systems
At Zzazz, led the recovery of MongoDB systems after critical indexes were deleted during a security incident. Rebuilt database indexes, restored application stability, coordinated recovery under pressure, and helped the business regain access to affected systems.
- Restored database performance and application availability
- Owned hands-on debugging and recovery during a high-pressure incident
- Converted the incident into stronger operational safeguards
Cloud Migration
Led DigitalOcean to AWS Platform Migration
Headed the cloud-to-cloud migration from DigitalOcean to AWS at Zzazz, moving the platform toward stronger reliability, scalability, and long-term infrastructure control.
- Planned and executed migration work across cloud environments
- Improved platform scalability and operational maturity
- Aligned infrastructure decisions with reliability and cost goals
CloudWatch Logs Cost Optimization
Reduced CloudWatch Logs Spend by $10K per Month
CloudWatch Logs costs were rising because applications and platform components were sending high-volume logs into groups that did not all carry equal operational value. I investigated the ingestion pattern, separated useful signals from avoidable noise, and reduced recurring spend while preserving the logs needed for debugging and incident response.
- Reduced avoidable CloudWatch Logs ingestion
- Improved log group ownership and retention discipline
- Protected debugging and incident-response visibility
Read the case study
EKS Data Transfer Cost Optimization
Cut Inter-AZ and Cross-Region Traffic Costs by up to $25K per Month
AWS EKS clusters were generating expensive inter-AZ and cross-region data transfer charges. I analyzed service communication paths, workload placement, and traffic behavior to identify where network cost was being created unnecessarily.
- Reduced inter-AZ and cross-region AWS data transfer charges
- Improved visibility into EKS network cost drivers
- Optimized Kubernetes service and workload communication
Read the case study
AWS Direct Connect Cost Optimization
Learned to Right-Size AWS Direct Connect Capacity for Lower Recurring Cost
AWS Direct Connect is easy to treat as fixed infrastructure, especially when it supports production traffic. I learned to approach it as a FinOps and reliability decision: first understand real port utilization and traffic growth, then validate the redundancy design before changing capacity. That made it possible to reduce recurring Direct Connect cost without making availability a trade-off.
- Connected AWS billing data with port utilization and traffic trends
- Validated redundancy and failure scenarios before reducing capacity
- Right-sized Direct Connect for actual demand rather than historic assumptions
Read the case study
VMware Oracle Licensing Optimization
Reduced Oracle Licensing Costs on VMware Clusters
Oracle workloads running across broad VMware clusters can increase licensing exposure when placement is not tightly controlled. I reduced that exposure by segregating ESXi hosts dedicated to Oracle VMs into a separate cluster and applying host affinity rules so Oracle workloads stayed on approved capacity.
- Separated Oracle VMs from general-purpose VMware clusters
- Controlled placement with ESXi host affinity rules
- Reduced unnecessary licensing scope and cost exposure
Read the case study