Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

SpeechQuest

A professional speech-to-text and question generation application that operates entirely locally without relying on external APIs.

Features

  • Audio Transcription - Convert speech to text from:
    • Audio file uploads
    • YouTube URLs
    • Live microphone recording
  • Question Generation - Automatically generate relevant questions from transcribed content
  • Offline Processing - All processing is done locally using pre-trained models
  • Modern UI - Clean, responsive interface built with React and TailwindCSS

Architecture

Backend (Python/FastAPI)

  • Speech-to-Text: Uses Faster-Whisper, a highly optimized version of OpenAI's Whisper model
  • Question Generation: Implements the T5 transformer model for generating questions
  • RESTful API: FastAPI-based interface with automatic documentation
  • Media Processing: Support for various audio formats and YouTube URL extraction

Frontend (React/TailwindCSS)

  • Component-based UI: Modern React functional components with hooks
  • Responsive Design: Mobile-friendly layout using TailwindCSS
  • Asynchronous Processing: Progress indicators for long-running tasks
  • In-browser Recording: Direct microphone access via Web Audio API

Getting Started

Prerequisites

  • Python 3.8+
  • Node.js 14+
  • FFmpeg (for audio processing)

Installation

  1. Clone the repository
git clone https://github.com/yourusername/speechquest.git
cd speechquest
  1. Set up the backend
cd backend
pip install -r requirements.txt
  1. Set up the frontend
cd ../frontend
npm install

Running the Application

Option 1: Using the Start.bat File (Recommended)

Simply run the Start.bat file in the root directory and select your preferred option:

  1. Full Version - Starts both backend and frontend (requires Node.js)
  2. Simple Version - Starts backend and opens a simple HTML interface (no Node.js required)
  3. Backend API Only - Starts only the backend API service
  4. NLTK Resources Test - Verifies and downloads all required NLTK resources
  5. Question Generation Test - Tests if the T5 model and question generation system work properly
Start.bat

Option 2: Manual Startup

  1. Start the backend server
cd backend
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
# Download NLTK resources
python -c "import nltk; nltk.download('punkt', download_dir='C:/nltk_data'); nltk.download('punkt_tab', download_dir='C:/nltk_data'); nltk.download('stopwords', download_dir='C:/nltk_data'); nltk.download('averaged_perceptron_tagger', download_dir='C:/nltk_data'); nltk.download('wordnet', download_dir='C:/nltk_data')"
uvicorn app.main:app --reload
  1. Start the frontend development server
cd frontend
npm install
npm run dev
  1. Access the full application at http://localhost:5173 or the simple interface at file:///C:/NLP_Project/simple_interface.html

Technical Details

Speech-to-Text Process

The application uses Whisper, a neural speech recognition model, to transcribe audio. The implementation uses Faster-Whisper, which optimizes performance while maintaining high accuracy.

Question Generation Process

Question generation is handled by a fine-tuned T5 model that identifies key information in the text and creates relevant questions. The process involves:

  1. Text preprocessing and tokenization
  2. Keyword extraction to identify important topics
  3. Context-aware question generation using the transformer model
  4. Post-processing to ensure question quality and relevance

Troubleshooting

Question Generation Issues

If question generation is not working:

  1. Run the NLTK Resources Test from the Start.bat menu (Option 4) to verify all resources are properly installed
  2. Run the Question Generation Test from the Start.bat menu (Option 5) to test the T5 model and question generation system
  3. Check the Input Text Length - Text must be at least 10 words long to generate questions
  4. Restart the Application - Some resources may need to be reinitialized

API Connection Issues

If you see API connection errors:

  1. Check if the Backend is Running - Look for the terminal window showing the backend API
  2. Verify the API Health - Open http://127.0.0.1:8000/health in your browser
  3. Check for Port Conflicts - Make sure nothing else is using port 8000

General Issues

  1. Use Start.bat - Always use Start.bat to launch the application to ensure proper setup
  2. Check the Logs - Look for errors in the PowerShell window or in backend/speechquest.log
  3. Try the Simple Version - If the full version isn't working, try the simple version which has fewer dependencies

License

MIT

Acknowledgements

  • OpenAI Whisper model
  • Hugging Face Transformers library
  • The FastAPI framework
  • React and TailwindCSS communities

About

SpeechQuest — fully offline speech-to-text & automatic question generation with local models, no external APIs

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors