Data & AI Engineer | Cloud Data Platforms • Applied ML/NLP • RAG
I build reliable cloud data pipelines, analytics-ready datasets, machine learning systems, and source-grounded AI applications using Python, SQL, Azure, AWS, Databricks, PySpark, dbt, PostgreSQL/pgvector, scikit-learn, PyTorch, and Transformers.
I’m Rihua Van Steenburgh, a Data and AI Engineer focused on building reliable cloud data platforms, analytics-ready datasets, applied machine learning systems, and source-grounded RAG applications.
I am pursuing a Master of Science in Information Technology with a Data Analytics concentration at Middle Georgia State University, expected in December 2026. I also hold an Associate of Applied Science in Website Design/Development, which helps me connect data and AI backends with practical user-facing applications.
My portfolio combines cloud data engineering, analytics engineering, NLP model evaluation, and Generative AI application development. I focus on reproducible workflows, honest evaluation, testing, documentation, and clear technical communication.
-
Cloud Data Engineering: API ingestion, Azure Data Factory, ADLS Gen2, Databricks, PySpark, Delta Lake, AWS S3, ECS/Fargate, Redshift Serverless, and dbt
-
Applied Machine Learning & NLP: scikit-learn, TF-IDF, Linear SVM, PyTorch, Transformers, DistilBERT, leakage-safe evaluation, and error analysis
-
Generative AI & RAG: document ingestion, chunking, embeddings, PostgreSQL/pgvector, vector retrieval, source-grounded answers, citations, and safe no-answer behavior
-
Analytics Engineering & Reliability: SQL, dimensional modeling, analytics marts, data validation, pytest, GitHub Actions, Docker, evaluation, runbooks, and technical documentation
-
Languages: Python, SQL, R, JavaScript, TypeScript
-
Cloud & Data Platforms: Azure Data Factory, ADLS Gen2, Databricks, AWS S3, Redshift Serverless, ECS/Fargate, EventBridge, CloudWatch
-
Processing & Modeling: PySpark, Delta Lake, dbt, PostgreSQL, dimensional modeling, star schema
-
AI / RAG: document ingestion, chunking, embeddings, PostgreSQL/pgvector, vector search, retrieval-augmented generation, cited answers
-
Orchestration & CI/CD: Airflow, GitHub Actions, workflow automation, validation checks
-
BI & Tools: Power BI, DAX, Docker, Git/GitHub, Jupyter Notebook, VS Code
-
Web & APIs: REST APIs, JSON, CSV, WordPress
CivicLens RAG — NYC 311 Operations Copilot ( https://github.com/rihua-tech/civiclens-rag-nyc311 )
Hybrid RAG application that ingests curated NYC 311 documentation, stores embeddings in PostgreSQL/pgvector, retrieves cited context, routes approved analytics questions, and presents grounded answers through a Streamlit interface.
Tech: Python, PostgreSQL, pgvector, embeddings, vector search, RAG, Streamlit, Docker, pytest, GitHub Actions
Financial Complaint Auto-Routing with NLP ( https://github.com/rihua-tech/financial-complaint-auto-routing-nlp )
Leakage-safe eight-class CFPB complaint-routing study comparing a TF-IDF and Linear SVM benchmark with a frozen DistilBERT challenger. The project includes group-aware splits, model evaluation, selective routing, and Human Review policies.
Tech: Python, scikit-learn, TF-IDF, Linear SVM, PyTorch, Transformers, DistilBERT, model evaluation
NYC 311 Service Requests Lakehouse (https://github.com/rihua-tech/nyc-311-service-requests-lakehouse)
Azure lakehouse pipeline using Azure Data Factory, ADLS Gen2, Databricks, PySpark, SQL, and Delta Lake to produce validated Bronze, Silver, Gold, fact, dimension, and analytics-mart outputs.
Cloud Flight Fare Pipeline (https://github.com/rihua-tech/cloud-flight-fare-pipeline)
AWS batch data pipeline using Docker, ECS/Fargate, EventBridge, S3, Redshift Serverless, SQL, and dbt with data-quality tests, CI checks, runbooks, and cloud execution proof.
I am upgrading CivicLens from a local Hybrid RAG prototype into a more complete AI application with:
- a versioned FastAPI backend;
- an optional real LLM provider;
- citation validation and safe provider-error handling;
- repeatable RAG and LLM evaluation reports;
- query, retrieval, and feedback logging;
- Dockerized Streamlit, API, and PostgreSQL/pgvector services;
- optional bounded agent and analytics-tool routing.
Thanks for stopping by! ✨


