a1 · AI4Science Gym — landscape

AI for Science // the LLM & agent tooling landscape

An exhaustive, filterable survey of the tools that plug an LLM/agent into the research workflowMCP servers, agent frameworks, lit-review & discovery agents, packaged skills, hosted products, scientific foundation models, and the benchmarks that grade them. For open repos & MCPs the lens is capabilities — every tool lists the discrete skills it exposes (read the skills column). For closed hosted products those same chips capture what researchers pay for — a demand signal. Built from a parallel research sweep; ?/license gaps are honest, not assumed.

tools cataloged
MCP servers
open-source
hosted products
distinct capabilities
Tool ▲▼ Type ▲▼ Domain ▲▼ Stage ▲▼ Skills / capabilities Open ▲▼ Self-host ▲▼ Access ▲▼ Stars ▲▼ Status ▲▼ Notes
Open open OSS code/weights · partial mixed/gated · closed proprietary Skills = the discrete tools/capabilities a repo exposes, or the headline features a hosted product sells. Stars = GitHub stars (—, n/a for hosted products).

How to read this landscape

Tools are tagged by where they sit in the research loop. Filter by Stage to assemble a stack end-to-end:

Discover → Read/Extract → Hypothesize → Experiment/Compute → Analyze → Write/Cite
A · MCP servers

The hands of the agent

Wrap a database, API, or compute backend behind the Model Context Protocol so any agent can call it. The biggest, fastest-moving bucket — literature (arXiv/PubMed/Semantic Scholar), biomedicine (the Augmented-Nature + Longevity-Genie + bio-mcp families), chemistry/materials, and data/compute (Jupyter, pandas, Slurm). The skills chips are the actual MCP tool names.

B · Agents & frameworks

The loops

From full autonomous "AI Scientist" loops (Sakana, RD-Agent, Agent Laboratory, the co-scientist reimplementation family) to lit-review/deep-research agents (PaperQA2, STORM, gpt-researcher), hypothesis/discovery engines (SciAgents, symbolic regression, causal discovery), and data/bioinformatics agents (Biomni, MetaGPT Data Interpreter). Most are research-grade — check Status and Stars.

C · Skills & hosted products

Packaged capability vs. demand signal

Packaged skills (Anthropic's doc skills, the K-Dense/AlterLab scientific skill-packs, research GPTs) are the open, composable form. Hosted products (Elicit, Consensus, Undermind, FutureHouse, Edison/Kosmos, Lila) are closed — but their feature chips tell you what researchers actually pay for: semantic search, claim/citation verification, structured extraction, chat-with-PDF, autonomous experiment design.

D · Models & benchmarks

Substrate and scoreboard

Domain foundation models (AlphaFold3/Boltz, ESM, RFdiffusion, MatterGen, GraphCast, Evo2) are the appendix — callable as tools, with license gotchas flagged (many weights are non-commercial even when code is open). Benchmarks (ScienceAgentBench, SciCode, CORE-Bench, LAB-Bench, MLE-bench, PaperBench…) are how the field grades agents; meta-lists are the awesome-lists to backfill from.

Coverage note (honest): compiled 2026-06-25 from a parallel multi-agent web sweep; star counts and licenses were read off live repo pages but are approximate as of that date, and ? means a LICENSE was genuinely not stated, not assumed. The biomedical MCP bucket is deepest (it is the most active corner of the ecosystem) and a handful of community star counts on very large skill-pack repos are fetch-reported, so treat the headline numbers as directional. Hosted-product internals are unknowable — their skill chips are advertised features, the demand signal. Deduped by repo across overlapping sweeps; a tool that is both a framework and a lit-review agent is filed once under its primary type.