Feynman: The AI Research Agent That Actually Understands Scientific APIs
Hook
Most AI research agents hallucinate citations or scrape PubMed HTML like amateurs. Feynman ships normalized connectors for 40+ scientific APIs—OpenAlex, PubMed, ChEMBL, gnomAD, GTEx—and treats literature review as a multi-agent coordination problem with auditable evidence chains.
Context
If you've ever asked an LLM to 'find papers on CRISPR off-target effects' or 'summarize the consensus on metformin and aging,' you've experienced the citation hallucination problem. GPT-4 will confidently cite papers that don't exist. Web search agents scrape PubMed result pages and miss full-text sections in Europe PMC. General-purpose autonomous agents like AutoGPT lack domain knowledge—they don't know that PMID↔DOI conversions require exact API calls, that citation graphs need OpenAlex's indexed relationships (not regex-parsed references), or that genomic variant lookups belong in gnomAD before you check clinical significance in ClinVar.
Feynman is an open-source TypeScript CLI that treats scientific research as a multi-agent workflow with specialized connectors. It's built atop the Pi agent runtime framework, orchestrating four agent personas (Researcher, Reviewer, Writer, Verifier) across literature search, citation analysis, paper ranking, and replication planning. The repository's real innovation is Feynman Bio Tools: a unified abstraction layer over 40+ scientific APIs spanning literature (OpenAlex, PubMed, Europe PMC, Crossref), genomics (Ensembl, gnomAD, GTEx, ENCODE), proteomics (STRING, UniProt), drug discovery (ChEMBL, PubChem, DrugBank), and clinical trials (cBioPortal, ClinicalTrials.gov). Instead of teaching an LLM to scrape, Feynman ships exact workflows that know how to traverse citation graphs, fetch section-aware full text, resolve identifiers across databases, and score papers on reproducibility dimensions before ranking them.
Technical Insight
Feynman's architecture separates three concerns: agent orchestration (the CLI and Pi runtime), skill modules (bundled under skills/ and prompts/), and the Bio Tools connector catalog. The standalone installer ships a self-contained native bundle with pinned Node.js to avoid dependency version drift—critical when your citation pipeline breaks because a transitive dependency updated OpenAlex query syntax. Authentication state lives in ~/.ahub for alphaXiv OAuth, while workbench sessions and installed Pi packages persist under ~/.feynman/orgs/<org_uuid>.
The PaperRank workflow is where Feynman's design gets interesting. It's not cosine similarity over embeddings—it's an iterative refinement pipeline that builds a local citation graph, scores papers on method reproducibility and provenance transparency, then optionally expands the graph and rescores with full-text evidence. Here's the execution model:
// Simplified PaperRank workflow (conceptual, not actual repo code)
async function rankPapers(query: string, options: RankOptions) {
// Step 1: Fetch initial candidate papers via OpenAlex/PubMed
const candidates = await bioTools.search(query, {
sources: ['openalex', 'pubmed', 'europepmc'],
limit: options.initialPool || 100
});
// Step 2: Build citation graph (cited-by and references)
const graph = await buildCitationGraph(candidates);
// Step 3: Score each paper on reproducibility dimensions
const scored = candidates.map(paper => ({
...paper,
score: calculateScore({
citationCount: graph.inDegree(paper.id),
methodClarity: analyzeMethodsSection(paper.abstract),
dataAvailability: checkDataDeposits(paper.links),
codeAvailability: checkGitHubLinks(paper.fullText)
})
}));
// Step 4: Optional citation expansion
if (options.expandCitations) {
const expanded = await expandGraph(graph, options.expandCitations);
scored.push(...scoreExpandedPapers(expanded));
}
// Step 5: Optional full-text rescoring for top-N
if (options.fullTextTop) {
const topPapers = scored.slice(0, options.fullTextTop);
for (const paper of topPapers) {
const fullText = await bioTools.fetchFullText(paper.id);
paper.score = rescoreWithFullText(paper, fullText, {
checkMethods: true,
checkResults: true,
checkReproducibility: true
});
}
}
return scored.sort((a, b) => b.score - a.score);
}
The --expand-citations 2 flag walks the citation graph two hops out, pulling in papers that cite or are cited by your initial candidates. This is computationally expensive—a single influential paper might expand the pool by 500+ works—but it surfaces foundational methods papers and recent follow-ups that keyword search misses. The --full-text-top 10 flag fetches complete paper text (not just abstracts) for the top 10 ranked papers, extracts Methods/Results/Discussion sections using Europe PMC's XML API, then rescores based on reproducibility checklists: Does the Methods section specify exact software versions? Are datasets deposited in public repositories? Is analysis code linked? This section-aware analysis produces evidence chains—'Paper X ranks #1 because the Methods section specifies DESeq2 v1.34.0, raw data is in GEO accession GSE123456, and analysis code is at github.com/lab/project'—that you can audit.
The Bio Tools connector catalog is the real differentiator. Instead of teaching an LLM to parse API documentation, Feynman ships 40+ pre-built connectors with exact workflows. Need to check if a gene variant is pathogenic? The connector chain is gnomAD (population frequency) → ClinVar (clinical assertions) → OMIM (phenotype associations). Need drug-target interactions? ChEMBL (binding assays) → PubChem (compound properties) → DrugBank (approved indications). Need to find transcription factor binding sites? ENCODE (ChIP-seq peaks) → JASPAR (motif matrices) → UniBind (predicted sites). Each connector handles authentication, rate limiting, identifier resolution, and response normalization:
// Example Bio Tools connector usage (conceptual)
const variant = await bioTools.gnomad.lookup({
chromosome: '17',
position: 43044295,
ref: 'G',
alt: 'A'
});
const clinical = await bioTools.clinvar.lookup({
rsid: variant.rsid
});
const phenotypes = await bioTools.omim.lookup({
gene: variant.geneSymbol
});
return {
variant: variant.id,
populationFrequency: variant.alleleFrequency,
pathogenicity: clinical.clinicalSignificance,
associatedPhenotypes: phenotypes.map(p => p.name)
};
The multi-agent decomposition is workflow-specific. The /deepresearch command spawns parallel Researcher agents for source-heavy investigation—one agent might query OpenAlex for papers while another fetches preprints from bioRxiv and a third checks clinical trials registries. The /lit command runs a literature review workflow where multiple Reviewer agents independently assess papers, then a Writer agent synthesizes consensus findings and flags disagreements. The /audit command compares paper claims against linked codebases by spinning up a Verifier agent that clones repos and checks if analysis scripts match Methods sections. The /replicate command plans replication workflows—Docker containers, compute environments, data dependencies—but gates execution behind explicit user approval. You see the plan, approve resource provisioning, then Feynman spins up Modal or RunPod compute to run the replication check.
The local workbench (feynman serve) is essentially a personal science IDE built on SQLite. It tracks artifact lineage (which query produced which paper set), manages compute host inventory (local/SSH/cloud contexts for replication workflows), and renders 15+ scientific file formats: KET/RXN chemistry, VCF variants, FASTA sequences, XLSX workbooks, Jupyter notebooks. You can annotate elements in HTML reports, attach notes to artifacts with modal previews, and export everything to cloud storage with audit logs. This isn't a web wrapper around chat—it's a stateful research environment where you can trace 'why did this paper rank here?' by inspecting the citation graph, full-text evidence, and scoring rubric.
Gotcha
The 40+ API connectors create a maintenance nightmare. Feynman depends on stable APIs from OpenAlex, PubMed, ChEMBL, ENCODE, gnomAD, GTEx, and dozens more. When any of these services change response schemas, deprecate endpoints, or introduce rate limits, your research workflows silently degrade. The repository shows no health checks or connector monitoring—you won't know that gnomAD lookups are returning stale data or that Europe PMC full-text extraction broke until you manually inspect results. This is inherent to the abstraction-over-40-APIs design: you're trading 'write your own API calls' complexity for 'trust that maintainers keep connectors updated' risk.
PaperRank's citation expansion and full-text scoring don't expose cost controls. Running --expand-citations 2 --full-text-top 10 on an influential paper could transitively fetch 500+ papers, make hundreds of API calls, and send 50-page PDFs to your LLM for section extraction. There's no progress UI, no incremental checkpointing, and no budget caps—the operation runs until completion or failure. If you're using a paid LLM backend, you might burn through $50 of API credits before realizing the query scope exploded. The Bio Tools connectors are biologically skewed: genomics, proteomics, drug discovery are well-covered, but physics (no crystallography or spectroscopy APIs), chemistry (only basic PubChem support, no reaction databases), and social sciences (no survey data or census APIs) get minimal tooling. If you're researching materials science or economics, Feynman's connector catalog won't help.
Verdict
Use Feynman if you're doing biomedical or life sciences research where citation accuracy and reproducibility matter, you're comfortable adopting a CLI workflow, and you need the 40+ scientific API connectors that no general-purpose agent ships. The PaperRank scoring with citation graph expansion and full-text analysis delivers reproducible, auditable literature reviews that AutoGPT-style tools can't match. The local workbench with artifact lineage and compute orchestration is a genuine research environment. The skills-only installer lets you extract the connector catalog into existing agent frameworks if you don't want the full CLI. Skip Feynman if you're outside biomedical domains (physics, chemistry, social science tooling is minimal), if you need connector health monitoring or multi-agent budget controls (no observability for degraded APIs or runaway costs), or if you can't accept the standalone installer's all-or-nothing upgrade model. For exploratory questions across all domains, Perplexity Pro or Elicit offer better UX without the maintenance burden. For systematic biomedical reviews where you need to audit 'why does this paper rank #3?', Feynman is the only open-source tool that ships the full stack.