Skip to content
View rihua-tech's full-sized avatar

Block or report rihua-tech

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rihua-tech/README.md

Hi, I'm Rihua! 👋

Data & AI Engineer | Cloud Data Platforms • Applied ML/NLP • RAG

I build reliable cloud data pipelines, analytics-ready datasets, machine learning systems, and source-grounded AI applications using Python, SQL, Azure, AWS, Databricks, PySpark, dbt, PostgreSQL/pgvector, scikit-learn, PyTorch, and Transformers.


👨‍💻 About Me

I’m Rihua Van Steenburgh, a Data and AI Engineer focused on building reliable cloud data platforms, analytics-ready datasets, applied machine learning systems, and source-grounded RAG applications.

I am pursuing a Master of Science in Information Technology with a Data Analytics concentration at Middle Georgia State University, expected in December 2026. I also hold an Associate of Applied Science in Website Design/Development, which helps me connect data and AI backends with practical user-facing applications.

My portfolio combines cloud data engineering, analytics engineering, NLP model evaluation, and Generative AI application development. I focus on reproducible workflows, honest evaluation, testing, documentation, and clear technical communication.


🔎 What I Focus On

  • Cloud Data Engineering: API ingestion, Azure Data Factory, ADLS Gen2, Databricks, PySpark, Delta Lake, AWS S3, ECS/Fargate, Redshift Serverless, and dbt

  • Applied Machine Learning & NLP: scikit-learn, TF-IDF, Linear SVM, PyTorch, Transformers, DistilBERT, leakage-safe evaluation, and error analysis

  • Generative AI & RAG: document ingestion, chunking, embeddings, PostgreSQL/pgvector, vector retrieval, source-grounded answers, citations, and safe no-answer behavior

  • Analytics Engineering & Reliability: SQL, dimensional modeling, analytics marts, data validation, pytest, GitHub Actions, Docker, evaluation, runbooks, and technical documentation


🧰 Tech Stack

  • Languages: Python, SQL, R, JavaScript, TypeScript

  • Cloud & Data Platforms: Azure Data Factory, ADLS Gen2, Databricks, AWS S3, Redshift Serverless, ECS/Fargate, EventBridge, CloudWatch

  • Processing & Modeling: PySpark, Delta Lake, dbt, PostgreSQL, dimensional modeling, star schema

  • AI / RAG: document ingestion, chunking, embeddings, PostgreSQL/pgvector, vector search, retrieval-augmented generation, cited answers

  • Orchestration & CI/CD: Airflow, GitHub Actions, workflow automation, validation checks

  • BI & Tools: Power BI, DAX, Docker, Git/GitHub, Jupyter Notebook, VS Code

  • Web & APIs: REST APIs, JSON, CSV, WordPress


🚀 Featured Projects

CivicLens RAG — NYC 311 Operations Copilot ( https://github.com/rihua-tech/civiclens-rag-nyc311 )

Hybrid RAG application that ingests curated NYC 311 documentation, stores embeddings in PostgreSQL/pgvector, retrieves cited context, routes approved analytics questions, and presents grounded answers through a Streamlit interface.

Tech: Python, PostgreSQL, pgvector, embeddings, vector search, RAG, Streamlit, Docker, pytest, GitHub Actions

Financial Complaint Auto-Routing with NLP ( https://github.com/rihua-tech/financial-complaint-auto-routing-nlp )

Leakage-safe eight-class CFPB complaint-routing study comparing a TF-IDF and Linear SVM benchmark with a frozen DistilBERT challenger. The project includes group-aware splits, model evaluation, selective routing, and Human Review policies.

Tech: Python, scikit-learn, TF-IDF, Linear SVM, PyTorch, Transformers, DistilBERT, model evaluation

Azure lakehouse pipeline using Azure Data Factory, ADLS Gen2, Databricks, PySpark, SQL, and Delta Lake to produce validated Bronze, Silver, Gold, fact, dimension, and analytics-mart outputs.

AWS batch data pipeline using Docker, ECS/Fargate, EventBridge, S3, Redshift Serverless, SQL, and dbt with data-quality tests, CI checks, runbooks, and cloud execution proof.


🌱 Currently Building

I am upgrading CivicLens from a local Hybrid RAG prototype into a more complete AI application with:

  • a versioned FastAPI backend;
  • an optional real LLM provider;
  • citation validation and safe provider-error handling;
  • repeatable RAG and LLM evaluation reports;
  • query, retrieval, and feedback logging;
  • Dockerized Streamlit, API, and PostgreSQL/pgvector services;
  • optional bounded agent and analytics-tool routing.

📫 Connect With Me


Thanks for stopping by! ✨

Pinned Loading

  1. civiclens-rag-nyc311 civiclens-rag-nyc311 Public

    Hybrid RAG operations copilot with cited answers, PostgreSQL/pgvector retrieval, analytics routing, Streamlit, evaluation, Docker, and CI.

    Python

  2. financial-complaint-auto-routing-nlp financial-complaint-auto-routing-nlp Public

    Leakage-safe NLP decision-support study for CFPB complaint routing, comparing TF-IDF + Linear SVM with DistilBERT using selective routing, Human Review, and retrospective 2025 evaluation.

    Jupyter Notebook 1

  3. nyc-311-service-requests-lakehouse nyc-311-service-requests-lakehouse Public

    Azure medallion lakehouse for NYC 311 service-request analytics using ADF, ADLS Gen2, Databricks/PySpark, Delta Lake, data-quality checks, dimensional modeling, and Power BI-ready marts.

    Python

  4. cloud-flight-fare-pipeline cloud-flight-fare-pipeline Public

    End-to-end AWS batch data pipeline for flight-fare analytics using EventBridge Scheduler, ECS/Fargate, S3, Redshift Serverless, dbt marts and tests, Docker, and CloudWatch execution proof.

    Python

  5. data-engineer-portfolio data-engineer-portfolio Public

    Professional data engineering portfolio showcasing Azure and AWS cloud pipeline projects, analytics-ready datasets, and modern data tooling.

    TypeScript