An assistive keyboard for patients with communication disabilities, built with an NLP interface for common language interactions.
- Project Overview
- Application System Design Diagram
- Deployment
- Tech Stack
- Setup Instructions
- Usage
- Contributing
- License
This project aims to create an assistive keyboard to help patients with communication disabilities. It features a Natural Language Processing (NLP) interface that facilitates common language interactions, making it easier for users to communicate effectively. The project also integrates with WebGazer, a web-based gaze tracking library that enables users to interact with the interface using eye movements.
- NLP Interface: Utilizes advanced NLP techniques to predict and suggest common phrases, reducing the effort required for patients to communicate.
- Gaze Tracking: Integrates WebGazer to allow users to control the keyboard and select options using their eye movements, providing an alternative input method for those with limited mobility.
- Voice Input: Allows users to input text using their voice, providing another accessible input option.
- Customizable Interface: Includes options for high contrast and dark mode to accommodate different visual preferences and needs.
- Customizable Vocabulary: Allows caregivers and users to customize the vocabulary and phrases used by the NLP model to better suit individual needs.
- Cross-Platform Compatibility: Built with web technologies to ensure compatibility across different devices and operating systems.
- Secure and Private: Ensures user data and interactions are secure and private, adhering to best practices in data security.
The project is deployed and can be accessed at the following links:
Note: Since the applications are hosted on one of your group member's Linux VPS, you might experience increased latency with the Whisper CPP model, as it provides inference through CPU
We recommend cloning and running the project components locally in a build mode for the optimal experience.
The project is built using the following technologies:
- Frontend: JavaScript, ReactJS, TailwindCSS, Vite
- Backend: Python, FastAPI,
- Machine Learning: Whisper CPP model from HuggingFace and WebGazer.js
- Third-party LLM API: Gemini 2.5 Pro
- Other: Docker for containerization of application components
- Option 1: Docker and Docker Compose installed to run components in isolated container environment
- Option 2: Node, NPM, Python 3.11+, PIP, FFMPEG installed on machine
- You can log into Google AI Studio with your account, and gneerate a free API key here: https://aistudio.google.com/app/apikey
-
Clone the repository:
git clone https://github.com/sfu-cmpt340/2025_1_project_18.git cd 2025_1_project_18 -
Create two
.envfiles in/frontendand/backendfolders, relative to the root and follow thetemplate.envkey=value pairs:cd frontend touch .env cd ../backend touch .env
# backend values: GOOGLE_GENAI_API_KEY=your-gemini-api-key MODE=dev/production PORT=integer-port-value-free-on-machine FRONTEND_APP_URL=url-with-http-protocol WHISPER_MAXIMUM_THREADS_USAGE=integer-for-thread-usage WHISPER_MODEL_NAME=file-name-pretrained-weights-folder
# frontend values VITE_BACKEND_API_BASE_URL=url-with-http-protocol
-
Run the setup script:
chmod +x setup.sh ./setup.sh
-
Frontend Setup:
- Navigates to the
frontenddirectory. - Installs the necessary frontend dependencies using
npm install.
- Navigates to the
-
Backend Setup:
- Navigates to the
backenddirectory. - Creates a Python virtual environment.
- Activates the virtual environment.
- Installs the
whisper-cpp-pythonpackage with specific build options. - Installs other backend dependencies from the
requirements.txtwithpip install -r requirements.txt. - Deactivates the virtual environment.
- Navigates to the
-
Model Download:
- Downloads the specified Whisper model from HuggingFace and saves it in the
backend/pretrained_weightsdirectory.
- Downloads the specified Whisper model from HuggingFace and saves it in the
-
Note:
- If you don't want to use Docker, you can naviage to the respective directories for each component for the project, and follow the previous commands on each section
To run the project, you can use Docker Compose:
-
At the root level of the project, run:
docker compose up
-
To stop the project, run:
docker compose down
We welcome contributions from the community. If you would like to contribute, please follow these steps:
- Fork the repository.
- Create a new branch (
git checkout -b feature-branch). - Make your changes and commit them (
git commit -am 'Add new feature'). - Push to the branch (
git push origin feature-branch). - Create a new Pull Request.
This project is licensed under the MIT License. See the LICENSE file for more information.
- Whisper CPP Python binding
- Hugging face models for Whisper CPP
- FFMPEG Audio Resampling Docs
- Librosa Scientific Audio Processing Docs
- Text Streaming Article over FastAPI endpoint
- MediaRecorder class MDN official docs
- SpeechRecognition class MDC official docs
- WebGazerJS library
- Google Gemini API Text Generation Docs
- TextBlob NLP Library Docs
- Common Medical words dataset from Kaggle
- Top 10,000 Most Common words dataset sourced from Google
- Software to create system design application diagram
