Brief: A Knowledge-Base Approach to Detecting Project Toolchains in One Command
Hook
Most toolchain detection is hardcoded if-else chains that break when tools update. Brief flips this: 355 declarative TOML files that you can patch without recompiling, turning tooling archaeology from a maintenance nightmare into a data problem.
Context
Drop into an unfamiliar repository and you're immediately stuck: How do I run tests? What linter rules apply? Does this project use a framework with known security pitfalls? You could grep for pytest or golangci-lint configs, parse package.json scripts, or dig through CI configs—but that takes 20 minutes and fails the moment a project uses a tool you didn't anticipate.
This gap hits hardest in two scenarios. First, AI coding agents need bootstrapping context: an LLM can't propose a fix using black if it doesn't know black is configured, and it wastes tokens analyzing code in languages the project doesn't use. Second, security and supply-chain teams need to inventory toolchains across dozens of repositories without manually auditing each one. Existing solutions split the problem: linguist detects languages but ignores tooling, syft extracts dependencies but not linters, and custom scripts hardcode assumptions that break on the next repository. Brief consolidates this into a single-binary CLI that outputs structured JSON describing everything from test frameworks to dangerous function calls.
Technical Insight
Brief's architecture centers on a knowledge base of 355 tool definitions stored as TOML files in data/tools/. Each definition is a declarative signature that specifies how to recognize a tool without writing detection logic. Here's the definition for pytest:
name = "pytest"
slug = "pytest"
description = "Python testing framework"
homepage = "https://pytest.org"
tags = ["role:testing", "lang:python", "type:framework"]
[[files]]
path = "pytest.ini"
[[files]]
path = "pyproject.toml"
section = "tool.pytest.ini_options"
[[commands]]
name = "pytest"
The detection pipeline is a four-stage pattern matcher: filesystem traversal walks up to 8 directory levels by default, building an index of files and parsing manifest files like package.json or Cargo.toml with language-specific parsers. The matching engine then loads all tool definitions and checks each signature—pytest.ini exists, or pyproject.toml contains a [tool.pytest.ini_options] section, or pytest appears as a command in scripts. Tools that match get added to the output with their taxonomy tags intact.
This declarative approach has a killer advantage: extending coverage doesn't require code changes. Want to add support for a new Rust linter? Create a TOML file with file signatures and tags. The knowledge-base-as-code pattern means contributions don't need Go expertise—just understanding of the tool being added. Compare this to hardcoded detection where adding pytest support means finding the right switch statement, writing regex for config parsing, and hoping you didn't break detection for nose or unittest.
The diff mode shows Brief's design philosophy: make detection context-aware without sacrificing speed. When you run brief diff main, Brief shells out to git to get changed files, then reruns the full detection pipeline but filters output to tools relevant to what changed. If you modified a .go file, you get golangci-lint and Go test frameworks but not pytest. If you edited .github/workflows/ci.yml, you get CI-related tools. The implementation is elegant—it maps file extensions and paths to tool signatures, so the filtering logic is derived from the same TOML knowledge base rather than duplicated:
// Pseudocode approximation of diff filtering
changedFiles := getGitDiff(baseBranch)
allTools := detectTools(repoPath)
relevantTools := []Tool{}
for _, tool := range allTools {
for _, sig := range tool.Files {
if anyChanged(changedFiles, sig.Path, sig.Pattern) {
relevantTools = append(relevantTools, tool)
break
}
}
}
This matters enormously for AI agents hitting context windows. Instead of feeding an LLM the full toolchain of a monorepo, you give it only the 3 tools relevant to the Python service that changed—cutting token counts by 80% while preserving precision.
The sink definitions feature layers another dataset on top: 700+ dangerous function patterns across 17 languages, stored per-tool in the knowledge base. When Brief detects that a project uses Express.js, it automatically outputs sinks like eval(), child_process.exec(), and vm.runInNewContext() that are relevant to Node.js. This isn't SAST—it doesn't analyze whether your code actually calls these functions—but it gives reviewers and agents a project-specific grep cheat sheet. For a security team onboarding to a new codebase, this is gold: you immediately know what to git grep for without needing language expertise.
The threat model feature uses conjunctive tag matching to map toolchains to CWE categories. If Brief detects tags role:framework AND layer:backend, it flags CWEs related to SQL injection and authentication bypasses. If it sees role:template_engine AND lang:javascript, it flags XSS categories. The lookup is deterministic—a static TOML table maps tag combinations to threat vectors—which avoids false positives from naive keyword matching but also means it can't reason about actual risk. You get a uniform threat list whether your Express app handles user input or just serves static files.
Manifest parsing preserves semantics that matter for supply chain analysis. When Brief parses Cargo.toml, it marks dependencies as direct, but when it parses Cargo.lock, those become transitive. Most detection tools flatten everything into one list, losing the distinction between 'we chose this' and 'this came 5 levels deep.' Brief's output includes scope too—devDependencies vs dependencies in npm, [dev-dependencies] vs [dependencies] in Cargo—so you can filter out test-only libraries when analyzing production attack surface.
Gotcha
Brief's 355-tool knowledge base is both its strength and ceiling. If your project uses an in-house linter or a niche framework not in the dataset, Brief simply won't see it. Extending coverage requires writing TOML definitions, which is easier than coding detection logic but still manual work—someone needs to research the tool's config files, command names, and taxonomy tags. For teams with heavily customized toolchains, you'll spend time authoring definitions before Brief becomes useful.
Static signature matching can't distinguish version-specific behavior. Brief detects golangci-lint by finding .golangci.yml, but it can't tell if that config uses linters introduced in v1.50 versus v1.30. You'll get a report saying 'golangci-lint detected' even if the config references features that don't exist in the installed version, leading to false confidence. Similarly, the threat model produces uniform CWE lists regardless of how you actually use a tool—an Express app gets SQL injection warnings even if it never touches a database.
External enrichment creates operational dependencies. When you run brief enrich, it hits ecosyste.ms for dependency metadata, endoflife.date for EOL status, and OpenSSF Scorecard for security scores. If any API is down, slow, or rate-limited, your command blocks or fails. There's no mention of caching beyond a git clone cache, so repeated runs make repeated requests. For CI pipelines or air-gapped environments, this is a non-starter—you can skip enrichment, but then you lose supply chain visibility. The 10k file limit can also silently truncate scans of large monorepos; the JSON output includes a flag indicating truncation, but if you're using human-formatted output interactively, you might not notice Brief only scanned half your repository.
Verdict
Use if: You're building AI coding agents that need bootstrapping context about arbitrary repositories, you're onboarding developers or auditing toolchains across multiple projects and need structured metadata in one command, or you're doing supply chain visibility work where you need dependency directness and manifest scope preserved. Brief's diff-aware detection and sink enumeration are purpose-built for these workflows, and the declarative knowledge base means you can extend it without Go expertise. Skip if: You need deep SAST findings rather than surface-level enumeration (Brief tells you dangerous functions exist, not whether you call them), your projects use heavily customized or proprietary tooling not in the 355-tool dataset (you'll spend time writing TOML definitions), or you require offline operation and can't depend on external APIs. For air-gapped environments or version-precise analysis, Brief's external enrichment and static signature matching are deal-breakers.