A multi-agent research system: Planner → Literature Search → Synthesis → Evaluation, with a working Reflexion-style self-critique/revise loop, tool use over the Model Context Protocol, parallelized execution, and persistent cross-session memory. Calls ServeLLM — a self-hosted, OpenAI-compatible LLM serving platform also built as part of this project — as its inference backend.
No paid API is used anywhere: not OpenAI, not Anthropic, not any other paid
provider. The standard openai Python client is used unmodified, but pointed at
ServeLLM's self-hosted endpoint — "OpenAI-compatible" here means only that the HTTP
shape matches; every actual inference call runs on a locally-hosted open-weight model.
- A real, working Reflexion loop, not just a described pattern: synthesis is
critiqued, scored, and revised up to N times with a genuine stopping condition —
and the score's reliability is tracked explicitly (
score_source:"parsed"/"heuristic"/"default"), because a 1.1B model asked to self-score ignores the format often enough that trusting it silently would be dishonest. - Tool use over MCP, not hardcoded function calls: arXiv search and sandboxed code
execution are both served by a real MCP server (
tools/mcp_server.py) that agent code talks to only by tool name and JSON shape — the same interface it would use against a tool server in a different process or language entirely. - Generated code is actually executed, not just produced — in a real, verified
network- and PID-isolated sandbox (
unshare, unprivileged namespaces — no Docker needed), not just resource limits. A network-access attempt from inside the sandbox is confirmed to fail withNetwork is unreachable, tested directly rather than assumed; every result reportsnetwork_isolatedexplicitly and stdout/stderr/exit code — including when the generated code fails, which it sometimes does. - Genuine parallel execution, right-sized to the workload: literature search and paper summarization are fanned out with a thread pool, cutting wall-clock time roughly in half (15.5s → 7.6s, measured on an identical question/config pair) — chosen deliberately over Ray/Redis after establishing the actual bottleneck was I/O-bound HTTP fan-out, not distributed compute.
- Persistent memory across sessions, backed by a real standalone Postgres instance: every completed research session is saved, and new questions are checked against past ones via keyword-overlap recall before synthesis — verified live with sequential runs where a related question correctly recalled a prior session and an unrelated one correctly recalled nothing.
- Semantic reranking of search results before they ever reach the LLM: every
deduped arXiv result is scored by real embedding similarity to the question
(
fastembed/ONNX, CPU-only, no torch, no paid API) and the bottom ones are dropped — measured 0.82-0.87 cosine similarity for genuinely relevant papers vs. 0.50-0.55 for irrelevant ones pulled in by arXiv's plain keyword search, and each retained paper is printed with its actual score, not just an LLM's say-so. - Confidence is a first-class output field, not something buried in a nested
value a caller has to know to check —
low_confidence/confidence_notesay outright when a score wasn't real or when the model's own critique still wasn't satisfied by the time the retry budget ran out. - Untrusted external content is treated as untrusted — arXiv abstracts are
delimited (instructions-vs-data) and heuristically scanned for injection-like
patterns before reaching any prompt, since they can influence what the code agent
later generates and executes. A flagged paper is marked (
injection_flagged), not silently trusted or silently dropped. - Sixteen real bugs found and fixed through actual live testing against a running
LLM backend — not hypothetical edge cases. Three examples:
httpxsilently not following a redirect meant every arXiv search returned zero results with no error; an MCP SDK major-version rename (FastMCP→MCPServer) crashed on first import and had to be root-caused by introspecting the installed package directly;pip installitself silently skipped installing a real dependency into this project's own environment because a same-named package already existed in an unrelated per-user directory. Full list in docs/ROADMAP.md.
Client (CLI / REST API)
│
orchestrator/pipeline.py
│
┌──────────┬──────────┬──────────┬──────────┐
│ │ │ │ │
PlannerAgent LiteratureAgent SynthesisAgent EvaluationAgent
│ │ (parallel │ │
│ │ fan-out) │ │
│ ▼ │ │
│ MCP client ──► MCP server ──► arXiv API │
│ │ │ │
│ ▼ agents/reranker.py (fastembed/ONNX,
│ CPU-only) — drop low-relevance results
│ before they reach the LLM
│ │ │
│ (optional) ───┴──► CodeAgent ──► ExperimentAgent
│ │
│ sandboxed subprocess
│
└──────────────────┬─────────────────────────────┘
▼
LLMClient (openai client,
pointed at ServeLLM)
│
▼
ServeLLM (self-hosted vLLM backend)
MemoryStore (optional) ──► Postgres — recall before synthesis, save after
Python · FastAPI · the openai client (against a self-hosted backend) · MCP
(Model Context Protocol) · SQLAlchemy · PostgreSQL · concurrent.futures ·
fastembed (ONNX Runtime, CPU-only embeddings) ·
arXiv API · NumPy.
conda create -n researchagent python=3.11 -y
conda activate researchagent
pip install -r requirements.txt
# requires a running ServeLLM instance (see the ServeLLM repo) — confirm it's up:
curl http://<serveLLM-host>:18742/healthz
# CLI: run one research question end-to-end
python scripts/run_research.py --base-url http://<serveLLM-host>:18742 \
--question "How can Mamba be made more efficient for time series forecasting?"
# optional: also generate and execute a demonstration script
python scripts/run_research.py --base-url http://<serveLLM-host>:18742 \
--question "How can a moving average be computed efficiently for streaming data?" \
--with-code
# or as a REST API
cp .env.example .env
bash scripts/start_api.sh
curl -X POST http://localhost:18800/v1/research \
-H "Content-Type: application/json" \
-d '{"question": "What are the main approaches to speculative decoding for LLM inference?"}'
bash scripts/stop_api.shPersistent memory (optional, off by default):
conda create -n researchagent-db -c conda-forge postgresql -y
bash scripts/init_memory_db.sh
python scripts/run_research.py --base-url http://<serveLLM-host>:18742 \
--question "How can Mamba be made more efficient for long sequences?" \
--memory-url "postgresql+psycopg://researchagent@/researchagent?host=$HOME/.researchagent-pg/run&port=18801"- Deterministic fallbacks alongside every LLM-dependent step, not just better prompting. A 1.1B model's search queries drifted off-topic under "keep it short" instructions alone; the fix was anchoring every query with deterministic keyword extraction from the question itself, verified by comparing real before/after search results, not assumed to help.
- Tool execution kept in a real MCP server process, not inlined, so the agent layer has no import-time dependency on how a tool is implemented — verified by probing the actual wire format (JSON-RPC over stdio) rather than trusting the SDK's own documentation, which didn't match the installed version's real API.
- Right-sized infrastructure, chosen after profiling, not by default: a stdlib thread pool for I/O-bound fan-out instead of Ray; Postgres (not SQLite) specifically because cross-session persistence needs a server that outlives any one process — both decisions explained in docs/ROADMAP.md, including where the named-but-unused technology would have been the wrong fit.
- Explicit tracking of unreliable model output rather than silently trusting format compliance: self-critique scores, code-block extraction, and search-query cleaning all carry a fallback path and a marker of which path was actually used.
See docs/ARCHITECTURE.md for the full design rationale and docs/ROADMAP.md for the phase-by-phase build log — every bug found, how it was diagnosed, and how it was actually fixed, verified against a live backend at every step rather than assumed to work.
ServeLLM currently serves Qwen2.5-7B-Instruct ("general," upgraded from TinyLlama-1.1B after reviewing real output quality honestly — see docs/ROADMAP.md) and Qwen2.5-Coder-1.5B ("code"). This pipeline's prompts, parsers, and fallbacks are still deliberately built assuming any model might ignore formatting instructions rather than trusting it will comply — that discipline caught real bugs even after the upgrade, and costs nothing when the model does comply. Swapping models requires no code changes here — the pipeline has no model-specific assumptions baked in.
This is a working prototype proving the architecture, not a system ready for
public/unauthenticated multi-user traffic. Scoped, agent-specific security hardening
is done (see the reranking/sandboxing highlights above); deliberately not done,
and why: scalability (job queue, multiple ServeLLM replicas), a managed/replicated
Postgres with backups, and a custom web UI (FastAPI's own /docs covers interactive
API access for free — see docs/ROADMAP.md's accessibility section for why a bespoke
frontend was skipped) would all be real, valuable engineering, but generic
production/systems work rather than agentic-AI work — see
ServeLLM for that side of this portfolio
instead of duplicating it here.