Skip to content
View EmilianFC20's full-sized avatar
💭
Creating
💭
Creating

Block or report EmilianFC20

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
EmilianFC20/README.md

Emiliano Fernández Cervantes

M.S. Computer Engineering at USC. I build hardware that runs machine learning, and I measure it. RTL and FPGA work, GPU kernels, and the systems glue in between.

Right now I'm building an INT8 attention accelerator on a Zynq-7000: a cycle-accurate C++ model, Verilog RTL co-simulated against it, and a PyTorch custom op on the board's ARM cores, so the final number is predicted vs. measured on real silicon, not a simulation result.


Selected work

Project What it is Result
zynq-attention-accelerator INT8 attention on a Zynq XC7Z020: C++ model → RTL → PyTorch op, each layer verified against the one above Phase 0 · no numbers published until measured
EE533-DPU Team DPU on NetFPGA: custom ARM-compatible CPU + from-scratch SIMD GPU, BF16 systolic tensor core, DMA engine Synthesized and run on Virtex-II Pro
netfpga-network-processor My EE 533 lab sequence: Verilog, C, CUDA, on real NetFPGA hardware Tiled CUDA matmul 72.83 ms vs 824.83 s CPU at N=4096
AEGIS (demo video) Hardware-accelerated intrusion-detection SmartNIC for power-grid infrastructure 850 ns ANN inference, measured on the board
family-assistant Self-hosted LLM stack: Ollama + Open WebUI + a RAG Discord bot, on a machine I built Benchmarked end to end (BENCHMARKS.md)
ansible-hospital-server-training Six progressive Ansible labs simulating a hospital server estate Inventory → hardening → roles → decommissioning

Also: INT4 quantization and Mixture-of-Experts fine-tuning on Llama 3.2-1B (USC EE 508), cutting model size 60.6%, benchmarked on an RTX 3070 Ti. Write-up pending publication.


Tools

Verilog · C/C++ · CUDA · PTX · Python · ARM assembly · Xilinx ISE · Cadence Virtuoso · ModelSim · Icarus Verilog · Docker · Ansible · Linux

Background

  • M.S. Computer Engineering, USC Viterbi. Viterbi Endowment Scholarship (full tuition)
  • Fulbright–García Robles grant (COMEXUS, 2024)
  • B.S. Biomedical Engineering, Tecnológico de Monterrey. Graduated top of the class

Website · LinkedIn

Pinned Loading

  1. zynq-attention-accelerator zynq-attention-accelerator Public

    INT8 attention accelerator on a Zynq-7000 SoC: cycle-accurate C++ model, Verilog RTL, and a PyTorch custom op, each verified against the layer above it.

    Python

  2. NetFPGA-Network-Processor NetFPGA-Network-Processor Public

    Verilog RTL, C, and CUDA for a NetFPGA-1G network processor: a custom ARM-compatible CPU paired with a from-scratch SIMD GPU, synthesized and run on Virtex-II Pro hardware (USC EE 533).

    C++

  3. USC-HW-Engineers/EE533-DPU USC-HW-Engineers/EE533-DPU Public

    The purpose of this project is to design a Data Processing Unit (DPU), integrating a custom-designed ARM-compatible processor with GPU acceleration.

    Verilog 1

  4. family-assistant family-assistant Public

    Discord bot bridging a family server to a self-hosted OpenWebUI/Ollama stack

    Python

  5. ansible-hospital-server-training ansible-hospital-server-training Public

    Hands-on Ansible training path simulating a hospital server environment: 6 progressive labs covering inventory, security, web portals, logging, roles, and decommissioning.

    Jinja

  6. Assembly_Simulator_ARM Assembly_Simulator_ARM Public

    Browser-based ARM assembly simulator, built to hand-verify test programs for a custom ARM-compatible CPU (USC EE 533).

    TypeScript