Autonomous Multi-Agent AI Platform for Scientific Research & Machine Learning Discovery
NeuroFlow AI (AutoML-Researcher) is a production-grade AI platform that acts as an autonomous Data Scientist. Rather than relying on a single monolithic prompt, NeuroFlow coordinates a network of specialized AI agents that dynamically source scientific datasets, engineer features, write machine learning pipelines, evaluate models, and compile professional research reports.
The platform provides a dual-interface experience: a lightweight CLI for terminal-based execution, and a state-of-the-art, high-fidelity Next.js Glassmorphism Dashboard that streams agent thought processes in real-time.
- Dynamic Agent Orchestration: Built on LangGraph state machines where distinct agents (Planner, Dataset, Coder, Evaluator, Writer) pass execution state autonomously.
- Human-in-the-Loop Safety Gate: Pauses workflow execution right before running machine learning scripts, allowing interactive human review and approval on the web UI.
- Real Scientific Data: Automatically interfaces with the OpenML API to fetch real-world datasets based on semantic natural language queries.
- End-to-End AutoML: Generates raw Python code to handle imputation, categorical encoding, train/test splitting, and algorithm fitting.
- Advanced Evaluation: Dynamically writes evaluation code to calculate accuracy, classification reports, and SHAP-based feature importance.
- Automated Research Reports: An AI Writer Agent takes the raw ML metrics and structures them into a professional Markdown-formatted research paper.
- Real-Time Streaming: A robust FastAPI backend streams agent execution logs (via WebSockets) to the frontend UI as the agents "think".
- Offline Mock Mode: Run full end-to-end testing (UI, WebSockets, CLI) without an active API key.
graph TD
A[User Query] --> B(Planner Agent)
B --> C(Dataset Agent)
C -->|Queries OpenML API| D(Coding Agent)
D -->|Writes Train Script| E(Evaluation Agent)
E -->|Writes Metrics/SHAP Code| F{Executor Engine}
F -->|Executes Code & Captures Output| G(Writer Agent)
G --> H[Final Markdown Report]
Backend
- FastAPI: Highly performant Python web framework for microservices and API gateways.
- LangGraph: Custom state-machine loop orchestrating the AI Agents.
- LangChain (Gemini Flash): Large Language Model core intelligence.
- Scikit-Learn & OpenML: Dataset fetching and algorithmic execution.
Frontend
- Next.js & React: Component architecture optimized for reactivity.
- Tailwind CSS & ShadCN: Ultra-fast utility CSS engine with beautiful accessible components.
- Framer Motion: Custom kinetic UI transitions and layout animations.
- WebSockets: Native browser WebSockets for real-time log streaming.
├── agents/ # Code for the specialized LangGraph agents
│ ├── base_agent.py
│ ├── planning_agent.py
│ ├── dataset_agent.py
│ ├── coding_agent.py
│ ├── evaluation_agent.py
│ ├── writer_agent.py
│ ├── state.py
│ └── graph.py # Custom pipeline orchestrator
├── backend/ # FastAPI app setup, schemas, and endpoints
│ ├── main.py
│ └── schemas.py
├── frontend/ # Premium Next.js + Tailwind Web Client UI
├── outputs/ # Output directories for generated reports/code
├── requirements.txt # Python dependencies manifest
└── run_pipeline.py # Root CLI pipeline interface
- Python 3.10+
- Node.js v18+ & npm v9+
- A Google Gemini API Key
- Clone the repository and navigate to the project root directory:
git clone https://github.com/yourusername/NeuroFlow-AI.git
cd "NeuroFlow-AI"- Initialize the Python Virtual Environment:
python -m venv .venv
source .venv/bin/activate- Install backend dependencies:
pip install -r requirements.txt- Install frontend dependencies:
cd frontend
npm install
cd ..Create a .env file in the root directory:
GOOGLE_API_KEY="your-gemini-api-key-here"Ensure your virtual environment is active (source .venv/bin/activate) before running python commands.
Start the backend API server with auto-reload:
uvicorn backend.main:app --reload --port 8000In a second terminal, start the Next.js frontend:
cd frontend
npm run devDashboard URL: http://localhost:3000
To run a research generation pipeline directly from your terminal:
python run_pipeline.py --query "Predict whether a patient has breast cancer"The CLI will stream the agent logs to the terminal and automatically save the raw Python code and final Markdown Research Report into the outputs/ folder!
MIT License