M.S. Computer Engineering at USC. I build hardware that runs machine learning, and I measure it. RTL and FPGA work, GPU kernels, and the systems glue in between.
Right now I'm building an INT8 attention accelerator on a Zynq-7000: a cycle-accurate C++ model, Verilog RTL co-simulated against it, and a PyTorch custom op on the board's ARM cores, so the final number is predicted vs. measured on real silicon, not a simulation result.
| Project | What it is | Result |
|---|---|---|
| zynq-attention-accelerator | INT8 attention on a Zynq XC7Z020: C++ model → RTL → PyTorch op, each layer verified against the one above | Phase 0 · no numbers published until measured |
| EE533-DPU | Team DPU on NetFPGA: custom ARM-compatible CPU + from-scratch SIMD GPU, BF16 systolic tensor core, DMA engine | Synthesized and run on Virtex-II Pro |
| netfpga-network-processor | My EE 533 lab sequence: Verilog, C, CUDA, on real NetFPGA hardware | Tiled CUDA matmul 72.83 ms vs 824.83 s CPU at N=4096 |
| AEGIS (demo video) | Hardware-accelerated intrusion-detection SmartNIC for power-grid infrastructure | 850 ns ANN inference, measured on the board |
| family-assistant | Self-hosted LLM stack: Ollama + Open WebUI + a RAG Discord bot, on a machine I built | Benchmarked end to end (BENCHMARKS.md) |
| ansible-hospital-server-training | Six progressive Ansible labs simulating a hospital server estate | Inventory → hardening → roles → decommissioning |
Also: INT4 quantization and Mixture-of-Experts fine-tuning on Llama 3.2-1B (USC EE 508), cutting model size 60.6%, benchmarked on an RTX 3070 Ti. Write-up pending publication.
Verilog · C/C++ · CUDA · PTX · Python · ARM assembly ·
Xilinx ISE · Cadence Virtuoso · ModelSim · Icarus Verilog · Docker · Ansible · Linux
- M.S. Computer Engineering, USC Viterbi. Viterbi Endowment Scholarship (full tuition)
- Fulbright–García Robles grant (COMEXUS, 2024)
- B.S. Biomedical Engineering, Tecnológico de Monterrey. Graduated top of the class


