Feynman: The Open-Source Research Agent That Ranks Papers Like a PI, Not an Algorithm
Hook
Most research agents rank papers by citation count and semantic similarity. Feynman scores them on reproducibility checklists, methods transparency, and provenance graphs—the signals your PI actually cares about when writing grants.
Context
Academic research tooling has bifurcated into two broken camps: generic LLM wrappers that hallucinate references and treat all papers as undifferentiated text blobs, and specialized science APIs (PubMed, ChEMBL, GTEx) that require bespoke integration work for every workflow. If you're a computational biologist trying to survey CRISPR literature, you'll spend a week stitching together OpenAlex for metadata, Europe PMC for full text, ChEMBL for compound lookups, and ClinVar for variant pathogenicity—then another week prompt-engineering GPT-4 to not fabricate DOIs. If you're an ML researcher auditing a NeurIPS paper's reproducibility claims, you're manually diffing GitHub repos against method sections with no structured lineage. SaaS tools like Elicit and Semantic Scholar solve paper search but lock your research memory in proprietary databases with no local state, no replication planning, and no transparency into why Paper A outranks Paper B.
Feynman treats research as a stateful, auditable workflow instead of stateless Q&A. Built by Companion Inc atop the Pi agent runtime, it's a TypeScript CLI that orchestrates specialized agents (Researcher, Reviewer, Writer, Verifier) across OpenAI, Anthropic, OpenRouter, and local LLMs, routing prompts to 40+ domain-specific APIs for chemistry, genomics, proteomics, and clinical trials. The architecture mirrors production agentic systems—tool-calling loops, SQLite-backed workbench state, cross-runtime skill distribution—but optimizes for scientific legibility over speed. PaperRank scoring surfaces reproducibility signals and citation graphs before synthesis. The /lit <lab> workflow maps publication trajectories for research groups. The /audit command diffs paper claims against codebases without auto-executing experiments. Everything lives in ~/.feynman/orgs/<org_uuid>/workbench as versioned frames, artifacts, and lineage records you can export, not ephemeral chat logs.
Technical Insight
Feynman's architecture distinguishes itself through three design choices that deviate from typical LangChain-style agent frameworks: transparent scoring primitives, distributed skill trees, and compute orchestration as instruction, not automation.
The PaperRank scoring system exposes what typical relevance models hide. Instead of embedding cosine similarity or opaque rerankers, Feynman chains explicit rubrics: citation count (with optional citation-graph expansion via OpenAlex), methods section transparency, reproducibility checklist adherence (does the paper link code? datasets? protocols?), and provenance scoring (preprint vs. peer-reviewed, journal impact, author H-index). When you run /deepresearch, parallel researcher agents fetch papers through a fallback cascade—OpenAlex → arXiv/AlphaXiv → DOI resolution → Europe PMC—then score each paper before synthesis. The scoring logic lives in packages/feynman/agents/researcher/tools/paperrank.ts and surfaces in reports as line-item justifications, not hidden weights. For grant writers and systematic reviewers, this transparency matters: you can defend why Paper X informed your background section when a program officer asks.
The skill distribution model decouples research primitives from the CLI wrapper. Feynman installs skills as file trees in .agents/skills/feynman, ~/.codex/skills/feynman, and .opencode/skills/feynman, letting Claude Projects, GitHub Codex, and OpenCode reuse paper-access workflows without adopting Feynman's REPL. A typical skill defines tool schemas and execution logic:
// Simplified skill structure from Feynman Bio Tools
export const pubchemSkill = {
name: 'pubchem_compound_lookup',
description: 'Fetch compound properties from PubChem by name, CID, or SMILES',
parameters: {
type: 'object',
properties: {
query: { type: 'string', description: 'Compound identifier' },
queryType: { enum: ['name', 'cid', 'smiles'] }
}
},
execute: async ({ query, queryType }) => {
const response = await fetch(
`https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/${queryType}/${encodeURIComponent(query)}/JSON`
);
const data = await response.json();
return {
cid: data.PC_Compounds[0].id.id.cid,
molecular_formula: data.PC_Compounds[0].props.find(p => p.urn.label === 'Molecular Formula').value.sval,
canonical_smiles: data.PC_Compounds[0].props.find(p => p.urn.label === 'SMILES').value.sval
};
}
};
This schema gets registered in Pi's tool catalog and surfaces to any agent runtime that imports the skill tree. A computational chemist using Claude Projects can invoke pubchem_compound_lookup without installing Feynman's CLI, and the skill's fetch logic stays synchronized across runtimes via Git submodules or npm packages. The 40+ connectors—OpenAlex, PubMed, ChEMBL, PubChem, Ensembl, GTEx, gnomAD, ClinVar, CADD, Rfam, CIViC, Open Targets, ENCODE—represent the most comprehensive open-science tool suite shipped with any OSS agent, but they're useful beyond Feynman's workflows.
Compute orchestration remains strictly instructional. When you run /replicate <paper>, Feynman generates Docker commands, Modal deployment scripts, or RunPod launch instructions but won't execute them. The agent plans the workflow—clone the repo, install dependencies, fetch datasets, modify hyperparameters—and formats it as reproducible bash or Python. This respects the boundary between research assistance and unsupervised code execution: you review the replication plan, adjust for your compute budget, then run it manually. The tradeoff is deliberate friction for safety. Auto-executing experiments in genomics or drug discovery risks burning compute budgets, violating data use agreements, or running unvetted code from papers. Feynman trusts you to hit Enter.
The workbench ledger treats research artifacts as first-class versioned entities. Every paper summary, replication plan, and lab trajectory gets stored in ~/.feynman/orgs/<org_uuid>/workbench/feynman-workbench.db as frames (conversation turns), artifacts (PDFs, LaTeX, Jupyter notebooks, chemistry KET/RXN structures, genome IGV tracks), lineage (which papers informed which syntheses), and credentials (OAuth tokens, API keys). The serve mode launches a local webapp with artifact previews—render LaTeX inline, visualize chemistry reactions with Ketcher, display genome tracks with IGV.js—so you're not juggling terminal output and browser tabs. The SQLite schema includes element-level HTML annotations: you can highlight a sentence in a paper frame, tag it as "contradicts Figure 3," and that annotation persists across sessions. This structured memory model beats chat transcripts when you're synthesizing evidence across dozens of papers over weeks.
Gotcha
Feynman's 40+ bio connectors lack batch-request optimization or aggressive caching beyond 'external fetched-content caching.' Each ChEMBL compound lookup, Ensembl VEP query, or GTEx tissue expression check likely hits upstream APIs individually. If you're processing 500 papers in a /deepresearch workflow and each spawns 10 tool calls, you're making 5,000 sequential HTTP requests. You'll hit rate limits on PubChem (5 requests/second without API keys), Europe PMC, and ClinVar. The architecture doesn't expose batch endpoints or request coalescing—the agents treat tools as stateless functions. For exploratory queries ("What do we know about BRCA1 variants?"), this works. For systematic reviews scraping 2,000 papers, you'll need manual caching middleware or risk IP bans.
The SQLite workbench backend creates a collaboration ceiling. All frames, artifacts, and lineage live in ~/.feynman/orgs/<org_uuid>/workbench/feynman-workbench.db with no documented sync, backup, or multi-user story. Two researchers on the same grant can't share a workbench without manually copying SQLite files or using Dropbox-style folder sync (which risks corruption). The cloud export audit logs suggest team features were planned but unshipped. If you're a solo PhD student, local-first SQLite is liberating—no vendor lock-in, no SaaS subscriptions. If you're a 10-person lab, you'll rebuild collaboration in Notion or Google Docs, losing the structured lineage and artifact versioning that make Feynman's workbench valuable. The README acknowledges organizations (~/.feynman/orgs/<org_uuid>) but doesn't explain multi-org workflows or role-based access.
Verdict
Use if: You're a computational biologist, ML researcher, or grant writer who needs transparent, auditable paper discovery with genomics/chemistry API access and you value local-first tooling over SaaS collaboration. The PaperRank scoring, /lit <lab> trajectory mapping, and 40+ bio connectors deliver legibility and domain coverage that generic ChatGPT wrappers cannot match. The replication planning workflows shine when you're auditing NeurIPS papers or surveying CRISPR literature and need citations you can defend in peer review. Skip if: You need collaborative workspaces (it's single-user SQLite with no sync), production-grade compute orchestration (it generates instructions, not automation), or minimal dependencies (the native bundle ships a full Node runtime). Also skip if you process thousands of papers per workflow (rate limits will hurt without batch optimization) or require telemetry-free operation (PostHog and OTLP are baked in). For exploratory solo research where reproducibility and provenance matter more than speed, Feynman is the most technically coherent open-source agent in the scientific tooling landscape.