Skip to content
View Abheenash's full-sized avatar

Highlights

  • Pro

Block or report Abheenash

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Abheenash/README.md

Hi, I'm Abheenash 👋

Typing SVG

Cloud · DevOps · Cloud Security · Terraform · Kubernetes · Serverless · GenAI · CI/CD · Kafka · Spark · C++ Systems

Website LinkedIn Live demo Houston, TX Open to work

Sixteen projects, all tested and running in CI. The thing worth knowing about them is that the numbers were measured on real hardware and real clusters — and the measurements that went against me are still in the READMEs.


Four times I was wrong

Skip the rest of this page if you like. This is the part that says what I'm actually like to work with.

I expected What happened What I did about it
Adding the obvious index would speed up the slowest check It made the run slower — the planner swapped one bulk hash join for 910,750 individual index probes Rewrote it as a window function over a sorted pass: 2,147 ms → 526 ms. Then measured index usage and deleted four indexes with zero scans (410 MB). →
My Kubernetes manifests were fine — every validator passed them Running them on a free local cluster found five bugs, including an alert meant to catch total silence that could never fire (in that language, "no data at all" is not zero) Fixed all five and wrote up the run, including the counter-test where killing every replica at once still cost ~2 s — because no manifest can prevent that. →
My five failure drills would trip five alarms Two never fired. One couldn't: the autoscaler replaced the instance faster than the alarm's window. The other's threshold was 80 connections on a database whose real ceiling is ~87 Both became different alarms — one on healthy host count, one at 70% of the instance class's actual limit. Validated in CI; not yet re-run live. →
Work-stealing would beat a simple global queue On heavy-tailed workloads it's a tie Published the tie. A benchmark table that only shows the rows where you won isn't a benchmark. →

Start here

aws-eks-platform — Kubernetes on AWS, the way a team would run it Terraform for EKS 1.35: Karpenter on Spot instead of fixed node groups, Pod Identity instead of IRSA, Argo CD so CI never holds cluster-admin, Prometheus with multi-window burn-rate alerts — validated, not applied. An earlier version ran live (66 resources, pod self-heal in 7 seconds, HPA 2 → 6 pods) and was then destroyed to avoid cost, so the current Kubernetes layer is proven on a free local cluster instead — a rolling deploy served 40/40 requests, a node drain 100/100, and force-killing every replica cost ~2 s, which is in the write-up too.

secure-container-pipeline — A build pipeline that actually stops things Four gates that fail the build, plus keyless signing on the way to ECR. PR #1 carried a deliberately planted AWS key and was blocked by two gates independently. Tests run against both a mocked DynamoDB and a real one. The container ships without a package installer — and the instruction that removes it now fails the build if it ever stops working, because it used to fail silently.

production-triage-toolkit — Fifteen read-only database diagnostics, in Java The problems support teams hear about from users — double bookings, ghost badges, stalled sync jobs, lock queues — each ranked, each with a runbook. 1,069 ms for all fifteen across 10 million rows. Read-only is verified with the server, not assumed. 176 tests, and a CI gate that fails the build if the README stops matching the code.

concurrent-kv-store — A Redis-shaped server in C++, on raw POSIX sockets A sharded store with TTLs, a 22-command protocol, append-only persistence, and three I/O models to compare: thread-per-connection, poll, and kqueue/epoll. 5.09 M requests/sec pipelined against 1.23 M with the original single global lock. Building it found five macOS/BSD socket bugs in itself — from POLLHUP handling to a listen backlog that dropped connections — each now pinned by a test. ThreadSanitizer- and AddressSanitizer-clean in CI.

The other twelve — click to expand

Operate

  • cloud-observability-sre — golden signals, SLOs with an error budget, burn-rate alerts, and a fault-injection experiment that uses the health alarm as its own stop condition. scripts/drill.sh throttles the function and times the page: 105 s.
  • aws-cloudops-lab — provision → break → detect → recover → document, five times on a real stack. Detection at 177 / 289 / 166 s; a timed database restore in 6 min 36 s against a 60-minute target.

Ship

  • aws-landing-zone — guardrails a workload can't undo: an account tree, five service control policies as plain JSON, deploy roles that trust only main of named repos. Each policy is run through a small permissions evaluator in tests — what it denies and what it must leave alone. Validated and scanned, not applied: creating member accounts can't be undone.

Build

  • serverless-file-share · live — files encrypted in the browser, the key living only in the URL fragment, so the backend stores something it cannot open. Tests now send real HTTP to the signed upload link with no credentials, and check that the storage service itself refuses a tampered signature.
  • job-hunt-command-center — the serverless tracker I run my own job search on. Generates a 2-page résumé from a job description via structured output, so it compiles and can't invent facts; classifies recruiter email through a queue-and-workflow pipeline.
  • portfolio-ai-assistant — the chat widget on my site. Every request logs its own cost, so the price is a graph: $0.0015 per answer. A 12-case prompt-injection eval is committed with its results.

Portable — the same service, three clouds

  • azure-container-platform · gcp-container-platform — the same API built again on Azure and on Google Cloud, so the three can be diffed. The write-up is the deliverable, not the app. Both are validated, not applied. One concrete difference, checked against all three real engines: a missing record makes one raise an error, one return an empty response, and one return a result object that says it doesn't exist.

Data & streaming

  • kafka-stream-processor — the three things that actually go wrong (a rebalance, a redelivery, one poison record), not the produce/consume every tutorial shows. The companion doc argues why my other project should stay on a simpler queue.
  • spark-lca-pipeline — a data pipeline rewritten for the size the data really is: millions of rows per year, one file per quarter, columns that drift between them. Runs locally, so this one has measured output rather than a plan.

Systems & parallel C++ — the rest of the C++ set, tested under thread and memory sanitizers

  • parallel-thread-pool — work-stealing, futures with exception propagation, a parallel loop whose waiting thread helps so nesting can't deadlock. 5.37× on 10 cores.
  • parallel-heat-diffusion — one grid, four backends, 320 combinations bitwise-identical to the serial reference. A memory-bandwidth probe sits beside the results and explains why 1.9× is the ceiling on a grid too big for cache, not a bug.

🛠️ Tech

AWS Amazon Bedrock Kubernetes / EKS Terraform Docker GitHub Actions Python Java PostgreSQL C++ Linux Bash

Cloud (AWS): Lambda · API Gateway · S3 · DynamoDB · Cognito · Bedrock · EventBridge · SQS · Step Functions · SES · SNS · EKS · ECS Fargate · ECR · VPC · ALB · CloudFront · Route 53 · KMS · Secrets Manager · IAM · WAF · CloudWatch · X-Ray · CloudTrail GenAI: Amazon Bedrock (Claude Sonnet 4.6 · Haiku · Opus) · AI résumé generation (structured JSON → LaTeX/PDF) · LLM classification & enrichment · JD↔résumé match scoring · grounded Q&A over a prompt-cached knowledge base (no vector store) · prompt-injection evals · per-request cost metering Containers & Kubernetes: Amazon EKS · Karpenter · EKS Pod Identity · AWS Load Balancer Controller · Argo CD (GitOps) · HPA · PodDisruptionBudgets · NetworkPolicy · Helm · kind · OpenShift port · ECS Fargate · Azure Container Apps · Google Cloud Run · Docker IaC & CI/CD: Terraform (AWS, azurerm, google) · terraform test with mocked providers · GitHub Actions · OIDC (keyless) · pre-commit · Renovate · tflint · Ansible · Jenkins Multi-cloud: Azure (Container Apps · Cosmos DB · ACR · managed identity) · GCP (Cloud Run · Firestore · Artifact Registry · workload identity federation) Data & streaming: Apache Kafka (MSK Serverless · Redpanda) · Apache Spark / PySpark (EMR Serverless) · DynamoDB · PostgreSQL DevSecOps & Security: IAM least privilege · Organizations SCPs & permissions boundaries · KMS/SSE encryption · Secrets Manager · WAF · Checkov · tfsec · Trivy · gitleaks · SBOM (CycloneDX) · cosign keyless signing · SLSA provenance · Dependabot Observability / SRE: CloudWatch dashboards & alarms · Prometheus & Grafana (kube-prometheus-stack, ServiceMonitor, PrometheusRule) · Datadog · Splunk · X-Ray · Synthetics · RUM · SLOs, error budgets & burn-rate alerts · anomaly detection · AWS Fault Injection Service · incident response · production triage & runbooks · PostgreSQL diagnostics Systems: POSIX sockets · epoll/kqueue · multithreading & work stealing · OpenMP · GDB · Valgrind · ThreadSanitizer & AddressSanitizer Languages: Python · Java · Bash · SQL · C++ · C · JavaScript · Go · Ruby · Perl Testing: pytest · moto · JUnit · testcontainers-style local services (LocalStack · DynamoDB Local · Firestore & Cosmos emulators · Redpanda · kind + Calico)


📜 Certifications

  • AWS Certified DevOps Engineer – Professional (DOP-C02) — verify
  • AWS Certified Solutions Architect – Associate (SAA-C03) — verify
  • AWS Certified Cloud Practitioner (CLF-C02) — verify

Top Languages

Full write-ups, with the war stories and the benchmarks behind every number above, at abheenash.com.

Pinned Loading

  1. cloud-observability-sre cloud-observability-sre Public

    Making a live serverless service observable + operable on AWS — dashboards, tracing, SLOs, alerting, incident-response demo

    HCL

  2. secure-container-pipeline secure-container-pipeline Public

    Hardened container service shipped through a security-gated DevSecOps pipeline on AWS (Fargate, Terraform, Trivy/Checkov/gitleaks)

    HCL

  3. aws-cloudops-lab aws-cloudops-lab Public

    Day-2 AWS operations lab — build, break, detect, recover, document. Incident drills, RCAs, restore tests, drift/import, patch compliance, cost.

    HCL

  4. aws-eks-platform aws-eks-platform Public

    Production-shaped Kubernetes app on Amazon EKS — Terraform, ALB Ingress, HPA autoscaling, keyless GitHub Actions CI/CD, and a pod-failure resilience drill.

    HCL

  5. job-hunt-command-center job-hunt-command-center Public

    Serverless job-search command center on AWS — track applications with the exact résumé sent, classify inbox replies, nudge stale applications. DynamoDB · Lambda · Cognito · S3 · EventBridge · Terra…

    Python