A professional speech-to-text and question generation application that operates entirely locally without relying on external APIs.
- Audio Transcription - Convert speech to text from:
- Audio file uploads
- YouTube URLs
- Live microphone recording
- Question Generation - Automatically generate relevant questions from transcribed content
- Offline Processing - All processing is done locally using pre-trained models
- Modern UI - Clean, responsive interface built with React and TailwindCSS
- Speech-to-Text: Uses Faster-Whisper, a highly optimized version of OpenAI's Whisper model
- Question Generation: Implements the T5 transformer model for generating questions
- RESTful API: FastAPI-based interface with automatic documentation
- Media Processing: Support for various audio formats and YouTube URL extraction
- Component-based UI: Modern React functional components with hooks
- Responsive Design: Mobile-friendly layout using TailwindCSS
- Asynchronous Processing: Progress indicators for long-running tasks
- In-browser Recording: Direct microphone access via Web Audio API
- Python 3.8+
- Node.js 14+
- FFmpeg (for audio processing)
- Clone the repository
git clone https://github.com/yourusername/speechquest.git
cd speechquest- Set up the backend
cd backend
pip install -r requirements.txt- Set up the frontend
cd ../frontend
npm installSimply run the Start.bat file in the root directory and select your preferred option:
- Full Version - Starts both backend and frontend (requires Node.js)
- Simple Version - Starts backend and opens a simple HTML interface (no Node.js required)
- Backend API Only - Starts only the backend API service
- NLTK Resources Test - Verifies and downloads all required NLTK resources
- Question Generation Test - Tests if the T5 model and question generation system work properly
Start.bat
- Start the backend server
cd backend
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
# Download NLTK resources
python -c "import nltk; nltk.download('punkt', download_dir='C:/nltk_data'); nltk.download('punkt_tab', download_dir='C:/nltk_data'); nltk.download('stopwords', download_dir='C:/nltk_data'); nltk.download('averaged_perceptron_tagger', download_dir='C:/nltk_data'); nltk.download('wordnet', download_dir='C:/nltk_data')"
uvicorn app.main:app --reload- Start the frontend development server
cd frontend
npm install
npm run dev- Access the full application at http://localhost:5173 or the simple interface at file:///C:/NLP_Project/simple_interface.html
The application uses Whisper, a neural speech recognition model, to transcribe audio. The implementation uses Faster-Whisper, which optimizes performance while maintaining high accuracy.
Question generation is handled by a fine-tuned T5 model that identifies key information in the text and creates relevant questions. The process involves:
- Text preprocessing and tokenization
- Keyword extraction to identify important topics
- Context-aware question generation using the transformer model
- Post-processing to ensure question quality and relevance
If question generation is not working:
- Run the NLTK Resources Test from the Start.bat menu (Option 4) to verify all resources are properly installed
- Run the Question Generation Test from the Start.bat menu (Option 5) to test the T5 model and question generation system
- Check the Input Text Length - Text must be at least 10 words long to generate questions
- Restart the Application - Some resources may need to be reinitialized
If you see API connection errors:
- Check if the Backend is Running - Look for the terminal window showing the backend API
- Verify the API Health - Open http://127.0.0.1:8000/health in your browser
- Check for Port Conflicts - Make sure nothing else is using port 8000
- Use Start.bat - Always use Start.bat to launch the application to ensure proper setup
- Check the Logs - Look for errors in the PowerShell window or in
backend/speechquest.log - Try the Simple Version - If the full version isn't working, try the simple version which has fewer dependencies
MIT
- OpenAI Whisper model
- Hugging Face Transformers library
- The FastAPI framework
- React and TailwindCSS communities