An exhaustive, filterable survey of the tools that plug an LLM/agent into the research workflow — MCP servers, agent frameworks, lit-review & discovery agents, packaged skills, hosted products, scientific foundation models, and the benchmarks that grade them. For open repos & MCPs the lens is capabilities — every tool lists the discrete skills it exposes (read the skills column). For closed hosted products those same chips capture what researchers pay for — a demand signal. Built from a parallel research sweep; ?/license gaps are honest, not assumed.
| Tool ▲▼ | Type ▲▼ | Domain ▲▼ | Stage ▲▼ | Skills / capabilities | Open ▲▼ | Self-host ▲▼ | Access ▲▼ | Stars ▲▼ | Status ▲▼ | Notes |
|---|
Tools are tagged by where they sit in the research loop. Filter by Stage to assemble a stack end-to-end:
Wrap a database, API, or compute backend behind the Model Context Protocol so any agent can call it. The biggest, fastest-moving bucket — literature (arXiv/PubMed/Semantic Scholar), biomedicine (the Augmented-Nature + Longevity-Genie + bio-mcp families), chemistry/materials, and data/compute (Jupyter, pandas, Slurm). The skills chips are the actual MCP tool names.
From full autonomous "AI Scientist" loops (Sakana, RD-Agent, Agent Laboratory, the co-scientist reimplementation family) to lit-review/deep-research agents (PaperQA2, STORM, gpt-researcher), hypothesis/discovery engines (SciAgents, symbolic regression, causal discovery), and data/bioinformatics agents (Biomni, MetaGPT Data Interpreter). Most are research-grade — check Status and Stars.
Packaged skills (Anthropic's doc skills, the K-Dense/AlterLab scientific skill-packs, research GPTs) are the open, composable form. Hosted products (Elicit, Consensus, Undermind, FutureHouse, Edison/Kosmos, Lila) are closed — but their feature chips tell you what researchers actually pay for: semantic search, claim/citation verification, structured extraction, chat-with-PDF, autonomous experiment design.
Domain foundation models (AlphaFold3/Boltz, ESM, RFdiffusion, MatterGen, GraphCast, Evo2) are the appendix — callable as tools, with license gotchas flagged (many weights are non-commercial even when code is open). Benchmarks (ScienceAgentBench, SciCode, CORE-Bench, LAB-Bench, MLE-bench, PaperBench…) are how the field grades agents; meta-lists are the awesome-lists to backfill from.