Skip to content
View qiaolinwang's full-sized avatar
  • New York

Highlights

  • Pro

Block or report qiaolinwang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
qiaolinwang/README.md

Hi, I'm Qiaolin (Chow-lin) Wang

Researcher at Bland · building audio general intelligence

  • 🔭 At Bland I work on expressive ASR, TTS, post-training, and speech-to-speech systems for voice AI in regulated industries.
  • 🎓 MS in Electrical Engineering, Columbia University (2026), where I was a research assistant on large audio language models with Prof. Nima Mesgarani. BE in Computer Science, Wuhan University (2024).
  • 🛠️ Before that: speech emotion recognition and multitask speech foundation models (Research Engineer intern), and TTS / voice cloning (ML Engineer intern) at WIZ.AI.
  • 🏠 Homepage: qiaolinwang.github.io · LinkedIn
  • ⚡ I'm also a hip-hop artist and producer. Find me on NetEase Cloud Music: Venti_J

📄 Publications

  • NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech, arXiv 2026 (submitted to ICASSP 2027). arXiv:2609.31892 · Audio samples
  • AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking, COLM 2026. arXiv:2601.17645
  • SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models, ICASSP 2026. IEEE · arXiv:2509.15661
  • Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations, EMNLP 2025 (SAC Highlight). arXiv:2509.15655

🧰 Tools

⭐ GitHub Stars in 3D

🎬 Where it started

My speech AI journey began as a side quest in 2022: teaching a model to speak as Paimon from Genshin Impact.

paimon_en_readme.mp4
  • Built and annotated a multi-speaker dataset of ≈48,000 clips (15 h) from 50 Genshin Impact characters, using ECAPA-TDNN for speaker classification and Whisper for transcription
  • Fine-tuned a VITS speech synthesis model on a curated set of Paimon clips
  • Released it as a technical demo on Bilibili (600K+ views) with a public Colab for anyone to try

Popular repositories Loading

  1. VITS VITS Public

    Forked from AlexandaJerry/vits-mandarin-biaobei

    Implementation of the VITS model

    Jupyter Notebook 398 74

  2. WHU_DB WHU_DB Public

    基于Pymysql和Pyqt的图书管理系统 Book Management System based on Pymysql and Pyqt

    Python 12 1

  3. TTS_Papers_Notes TTS_Papers_Notes Public

    2

  4. ChatVITS_Cyberpunk2077 ChatVITS_Cyberpunk2077 Public

    ChatGPT+VITS using Cyberpunk2077 dataset

    Python 2

  5. ESG_WEBUI ESG_WEBUI Public

    CITI BANK PROJECT DEMO

    HTML 2 4

  6. Text-to-sound-Synthesis Text-to-sound-Synthesis Public

    Forked from yangdongchao/Text-to-sound-Synthesis

    The source code of our paper "Diffsound: discrete diffusion model for text-to-sound generation"

    Python 1