Researcher at Bland · building audio general intelligence
- 🔭 At Bland I work on expressive ASR, TTS, post-training, and speech-to-speech systems for voice AI in regulated industries.
- 🎓 MS in Electrical Engineering, Columbia University (2026), where I was a research assistant on large audio language models with Prof. Nima Mesgarani. BE in Computer Science, Wuhan University (2024).
- 🛠️ Before that: speech emotion recognition and multitask speech foundation models (Research Engineer intern), and TTS / voice cloning (ML Engineer intern) at WIZ.AI.
- 🏠 Homepage: qiaolinwang.github.io · LinkedIn
- ⚡ I'm also a hip-hop artist and producer. Find me on NetEase Cloud Music: Venti_J
- NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech, arXiv 2026 (submitted to ICASSP 2027). arXiv:2609.31892 · Audio samples
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking, COLM 2026. arXiv:2601.17645
- SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models, ICASSP 2026. IEEE · arXiv:2509.15661
- Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations, EMNLP 2025 (SAC Highlight). arXiv:2509.15655
My speech AI journey began as a side quest in 2022: teaching a model to speak as Paimon from Genshin Impact.
paimon_en_readme.mp4
- Built and annotated a multi-speaker dataset of ≈48,000 clips (15 h) from 50 Genshin Impact characters, using ECAPA-TDNN for speaker classification and Whisper for transcription
- Fine-tuned a VITS speech synthesis model on a curated set of Paimon clips
- Released it as a technical demo on Bilibili (600K+ views) with a public Colab for anyone to try



