Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ResearchAgent

A multi-agent research system: Planner → Literature Search → Synthesis → Evaluation, with a working Reflexion-style self-critique/revise loop, tool use over the Model Context Protocol, parallelized execution, and persistent cross-session memory. Calls ServeLLM — a self-hosted, OpenAI-compatible LLM serving platform also built as part of this project — as its inference backend.

No paid API is used anywhere: not OpenAI, not Anthropic, not any other paid provider. The standard openai Python client is used unmodified, but pointed at ServeLLM's self-hosted endpoint — "OpenAI-compatible" here means only that the HTTP shape matches; every actual inference call runs on a locally-hosted open-weight model.

Highlights

  • A real, working Reflexion loop, not just a described pattern: synthesis is critiqued, scored, and revised up to N times with a genuine stopping condition — and the score's reliability is tracked explicitly (score_source: "parsed" / "heuristic" / "default"), because a 1.1B model asked to self-score ignores the format often enough that trusting it silently would be dishonest.
  • Tool use over MCP, not hardcoded function calls: arXiv search and sandboxed code execution are both served by a real MCP server (tools/mcp_server.py) that agent code talks to only by tool name and JSON shape — the same interface it would use against a tool server in a different process or language entirely.
  • Generated code is actually executed, not just produced — in a real, verified network- and PID-isolated sandbox (unshare, unprivileged namespaces — no Docker needed), not just resource limits. A network-access attempt from inside the sandbox is confirmed to fail with Network is unreachable, tested directly rather than assumed; every result reports network_isolated explicitly and stdout/stderr/exit code — including when the generated code fails, which it sometimes does.
  • Genuine parallel execution, right-sized to the workload: literature search and paper summarization are fanned out with a thread pool, cutting wall-clock time roughly in half (15.5s → 7.6s, measured on an identical question/config pair) — chosen deliberately over Ray/Redis after establishing the actual bottleneck was I/O-bound HTTP fan-out, not distributed compute.
  • Persistent memory across sessions, backed by a real standalone Postgres instance: every completed research session is saved, and new questions are checked against past ones via keyword-overlap recall before synthesis — verified live with sequential runs where a related question correctly recalled a prior session and an unrelated one correctly recalled nothing.
  • Semantic reranking of search results before they ever reach the LLM: every deduped arXiv result is scored by real embedding similarity to the question (fastembed/ONNX, CPU-only, no torch, no paid API) and the bottom ones are dropped — measured 0.82-0.87 cosine similarity for genuinely relevant papers vs. 0.50-0.55 for irrelevant ones pulled in by arXiv's plain keyword search, and each retained paper is printed with its actual score, not just an LLM's say-so.
  • Confidence is a first-class output field, not something buried in a nested value a caller has to know to check — low_confidence/confidence_note say outright when a score wasn't real or when the model's own critique still wasn't satisfied by the time the retry budget ran out.
  • Untrusted external content is treated as untrusted — arXiv abstracts are delimited (instructions-vs-data) and heuristically scanned for injection-like patterns before reaching any prompt, since they can influence what the code agent later generates and executes. A flagged paper is marked (injection_flagged), not silently trusted or silently dropped.
  • Sixteen real bugs found and fixed through actual live testing against a running LLM backend — not hypothetical edge cases. Three examples: httpx silently not following a redirect meant every arXiv search returned zero results with no error; an MCP SDK major-version rename (FastMCP → MCPServer) crashed on first import and had to be root-caused by introspecting the installed package directly; pip install itself silently skipped installing a real dependency into this project's own environment because a same-named package already existed in an unrelated per-user directory. Full list in docs/ROADMAP.md.

Architecture

                    Client (CLI / REST API)
                              │
                  orchestrator/pipeline.py
                              │
        ┌──────────┬──────────┬──────────┬──────────┐
        │          │          │          │          │
   PlannerAgent  LiteratureAgent  SynthesisAgent  EvaluationAgent
        │          │  (parallel      │               │
        │          │   fan-out)      │               │
        │          ▼                 │               │
        │    MCP client ──► MCP server ──► arXiv API  │
        │          │                    │               │
        │          ▼ agents/reranker.py (fastembed/ONNX,
        │            CPU-only) — drop low-relevance results
        │            before they reach the LLM
        │                        │                    │
        │          (optional) ───┴──► CodeAgent ──► ExperimentAgent
        │                                                  │
        │                                        sandboxed subprocess
        │
        └──────────────────┬─────────────────────────────┘
                            ▼
                   LLMClient (openai client,
                    pointed at ServeLLM)
                            │
                            ▼
              ServeLLM (self-hosted vLLM backend)

  MemoryStore (optional) ──► Postgres — recall before synthesis, save after

Tech stack

Python · FastAPI · the openai client (against a self-hosted backend) · MCP (Model Context Protocol) · SQLAlchemy · PostgreSQL · concurrent.futures · fastembed (ONNX Runtime, CPU-only embeddings) · arXiv API · NumPy.

Quickstart

conda create -n researchagent python=3.11 -y
conda activate researchagent
pip install -r requirements.txt

# requires a running ServeLLM instance (see the ServeLLM repo) — confirm it's up:
curl http://<serveLLM-host>:18742/healthz

# CLI: run one research question end-to-end
python scripts/run_research.py --base-url http://<serveLLM-host>:18742 \
  --question "How can Mamba be made more efficient for time series forecasting?"

# optional: also generate and execute a demonstration script
python scripts/run_research.py --base-url http://<serveLLM-host>:18742 \
  --question "How can a moving average be computed efficiently for streaming data?" \
  --with-code

# or as a REST API
cp .env.example .env
bash scripts/start_api.sh
curl -X POST http://localhost:18800/v1/research \
  -H "Content-Type: application/json" \
  -d '{"question": "What are the main approaches to speculative decoding for LLM inference?"}'
bash scripts/stop_api.sh

Persistent memory (optional, off by default):

conda create -n researchagent-db -c conda-forge postgresql -y
bash scripts/init_memory_db.sh

python scripts/run_research.py --base-url http://<serveLLM-host>:18742 \
  --question "How can Mamba be made more efficient for long sequences?" \
  --memory-url "postgresql+psycopg://researchagent@/researchagent?host=$HOME/.researchagent-pg/run&port=18801"

Notable engineering decisions

  • Deterministic fallbacks alongside every LLM-dependent step, not just better prompting. A 1.1B model's search queries drifted off-topic under "keep it short" instructions alone; the fix was anchoring every query with deterministic keyword extraction from the question itself, verified by comparing real before/after search results, not assumed to help.
  • Tool execution kept in a real MCP server process, not inlined, so the agent layer has no import-time dependency on how a tool is implemented — verified by probing the actual wire format (JSON-RPC over stdio) rather than trusting the SDK's own documentation, which didn't match the installed version's real API.
  • Right-sized infrastructure, chosen after profiling, not by default: a stdlib thread pool for I/O-bound fan-out instead of Ray; Postgres (not SQLite) specifically because cross-session persistence needs a server that outlives any one process — both decisions explained in docs/ROADMAP.md, including where the named-but-unused technology would have been the wrong fit.
  • Explicit tracking of unreliable model output rather than silently trusting format compliance: self-critique scores, code-block extraction, and search-query cleaning all carry a fallback path and a marker of which path was actually used.

See docs/ARCHITECTURE.md for the full design rationale and docs/ROADMAP.md for the phase-by-phase build log — every bug found, how it was diagnosed, and how it was actually fixed, verified against a live backend at every step rather than assumed to work.

Honesty about model capability

ServeLLM currently serves Qwen2.5-7B-Instruct ("general," upgraded from TinyLlama-1.1B after reviewing real output quality honestly — see docs/ROADMAP.md) and Qwen2.5-Coder-1.5B ("code"). This pipeline's prompts, parsers, and fallbacks are still deliberately built assuming any model might ignore formatting instructions rather than trusting it will comply — that discipline caught real bugs even after the upgrade, and costs nothing when the model does comply. Swapping models requires no code changes here — the pipeline has no model-specific assumptions baked in.

Path to production

This is a working prototype proving the architecture, not a system ready for public/unauthenticated multi-user traffic. Scoped, agent-specific security hardening is done (see the reranking/sandboxing highlights above); deliberately not done, and why: scalability (job queue, multiple ServeLLM replicas), a managed/replicated Postgres with backups, and a custom web UI (FastAPI's own /docs covers interactive API access for free — see docs/ROADMAP.md's accessibility section for why a bespoke frontend was skipped) would all be real, valuable engineering, but generic production/systems work rather than agentic-AI work — see ServeLLM for that side of this portfolio instead of duplicating it here.

License

MIT

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages