Nopus: A Deterministic Filter for LLM Verbosity in Coding Agents
Hook
Your coding agent just spent 400 tokens explaining 'paradigm shifts in capability governance frameworks' when you asked it to refactor a function. What if you could automatically catch and rewrite that nonsense before it clutters your conversation history?
Context
Coding agents have a verbosity problem. Ask Claude Code or Pi to explain a database schema, and you'll get three paragraphs of preamble about 'interesting considerations' and 'framework thinking' before any actual technical content. Unlike human code reviewers who adapt their explanations to context, LLMs default to lecture-hall framing because their training data includes everything from academic papers to Medium thought-leadership posts.
The standard solution—system prompts like 'be concise' or 'avoid jargon'—burns context window on meta-instructions that models inconsistently enforce. You're also stuck manually interrupting with '/rewrite that' commands when agents ignore your instructions, which requires vigilance every single response. Nopus takes a different approach: it analyzes completed agent responses using deterministic linguistic measurements, detects when prose crosses clarity thresholds, and automatically injects rewrite requests with specific evidence of what failed. No ML models, no prompt engineering, just precomputed frequency tables and concreteness ratings applied as a post-generation filter.
Technical Insight
Nopus operates as a platform-specific extension that hooks into the response stream after an agent completes generation. The architecture has three distinct phases: prose extraction, multi-metric analysis, and conditional rewrite injection.
The extraction phase strips code blocks, identifiers, and technical tokens to isolate natural language prose. It uses regex patterns to remove fenced code blocks, inline backticks, file paths, and variable names, leaving only the explanatory text agents use to frame their responses. This preprocessed prose then flows into the analysis pipeline.
The analysis engine runs seven deterministic measurements against packaged linguistic datasets. These aren't ML models—they're normalized lookup tables derived from peer-reviewed sources. SUBTLEX-US provides conversational word frequencies from film subtitles (representing spoken language patterns). Norvig web counts give internet-scale term distributions. Brysbaert concreteness ratings assign human-judged scores to 40,000 words on a 1-5 scale, where 'database' scores high but 'paradigm' scores low. The Carpentries technical glossary provides exemptions so terms like 'InteractiveSessionHost' don't trigger abstractness penalties.
Here's what the metric calculation looks like in simplified form:
interface ProseMetrics {
uncommonWords: number; // Words outside top 5000 SUBTLEX
veryUncommonWords: number; // Outside top 10000
abstractSentences: number; // Avg concreteness < 3.0
abstractVocabulary: number; // % abstract words (concreteness < 2.5)
nounStacks: number; // 3+ consecutive nouns
phraseDensity: number; // Multi-word prepositional phrases per sentence
formulaicCues: number; // 'Here's where', 'paradigm shift', etc.
}
function analyzeProseQuality(
text: string,
datasets: LinguisticDatasets,
sensitivity: 'low' | 'medium' | 'high'
): { passed: boolean; evidence: string[] } {
const metrics = computeMetrics(text, datasets);
const thresholds = THRESHOLD_PROFILES[sensitivity];
// Multi-metric threshold logic
const triggers: string[] = [];
if (metrics.uncommonWords > thresholds.uncommonWords) {
triggers.push(`Uncommon wording: ${extractExamples(text, 'uncommon')}`);
}
// Sustained abstraction: abstract vocabulary becomes critical
// when paired with high abstract sentence count
if (metrics.abstractSentences > thresholds.abstractSentences &&
metrics.abstractVocabulary > thresholds.abstractVocabulary) {
triggers.push(
`Sustained abstraction across ${metrics.abstractSentences} sentences: ` +
`${extractExamples(text, 'abstract')}`
);
}
if (metrics.formulaicCues > thresholds.formulaicCues) {
triggers.push(`Lecture-hall framing: ${extractExamples(text, 'formulaic')}`);
}
return {
passed: triggers.length === 0,
evidence: triggers
};
}
The threshold combinations are crucial. A single uncommon word doesn't trigger rewrite, but uncommon words paired with sustained abstraction (multiple sentences averaging concreteness < 3.0) do. This avoids false positives from technical terminology while catching genuine bloat like 'capability governance frameworks facilitate paradigm shifts in architectural thinking.'
When thresholds trigger, nopus doesn't just ask for a rewrite—it injects specific evidence into the conversation history:
function constructRewriteRequest(evidence: string[]): Message {
return {
role: 'user',
content: `That explanation needs clarification. Issues detected:\n\n` +
evidence.map(e => `- ${e}`).join('\n') + '\n\n' +
'Please rewrite focusing on concrete technical details without ' +
'preamble or abstract framing.'
};
}
This gives the LLM concrete learning signal—'here's where it gets interesting' is lecture-hall framing, 'governance framework' lacks concreteness—rather than vague 'be clearer' instructions. The single-retry bound is hardcoded: if the rewrite also fails thresholds, nopus accepts it anyway to prevent infinite loops.
Platform integration varies by agent architecture. Pi uses an extension API with transcript manipulation to hide rejected responses, so your conversation history shows only the successful rewrite. Claude Code and Codex use Stop hooks that intercept responses before they render. The datasets are preprocessed during nopus installation and validated with SHA-256 checksums, eliminating runtime downloads and ensuring reproducible analysis—same input always produces same verdict, making behavior debuggable unlike ML-based quality scoring.
Configuration lives in a simple JSON file at XDG_CONFIG_HOME:
{
"sensitivity": "medium",
"enabledPlatforms": ["pi", "claude-code"],
"technicalGlossary": [
"InteractiveSessionHost",
"WebAssembly",
"gRPC"
]
}
Users can add domain-specific technical terms to prevent false positives, and adjust sensitivity to control the uncommon-word and abstraction thresholds. The deterministic foundation means you can predict exactly when nopus will intervene once you understand your threshold configuration.
Gotcha
The English-only limitation is absolute. SUBTLEX, Norvig counts, and Brysbaert ratings are all English-centric datasets with no multilingual equivalents packaged. If your coding agent responds in Spanish, French, or code-switches between languages, nopus will produce garbage metrics. The technical glossary exemption only works for English technical terms.
Effectiveness validation is thin. The 5.3%-18.6% rewrite rate comes from analysis of Pi responses, but there's no published data showing whether rewrites actually improved task completion or developer comprehension. You're trusting that 'more concrete, less abstract' correlates with usefulness, which might not hold if you're asking conceptual questions where abstraction is appropriate. The tool has no feedback mechanism—it can't learn whether you found its interventions helpful or annoying, so calibration depends entirely on manually adjusting sensitivity settings until false-positive rates feel tolerable.
Code-stripping heuristics create a blind spot for prose embedded in strings, comments, or inline documentation. An agent could generate a perfect external explanation while writing incomprehensible code comments full of vague abstraction, and nopus would pass it. Conversely, legitimate use of abstract technical vocabulary in markdown explanations might trigger false positives if terms aren't in the Carpentries glossary.
Platform support is limited to Pi, Claude Code, and Codex. Popular alternatives like Cursor, Aider, Continue, or Cody aren't supported because each requires reverse-engineering platform-specific extension APIs. There's no generic proxy mode that could work across arbitrary coding agents by intercepting HTTP requests—you're locked into the platforms nopus explicitly integrates with.
Verdict
Use if: You're doing extended coding sessions with Pi, Claude Code, or Codex where agent verbosity actively slows you down, and you'd rather have automated intervention than manually type '/rewrite that' every third response. The deterministic thresholds make it safe to enable permanently without worrying about ML-based false positives spiraling into infinite rewrites. You value predictable behavior and can tolerate occasional false positives by tweaking sensitivity settings. Skip if: You work in non-English contexts, use unsupported platforms like Cursor or Aider, or you're in learning mode where verbose explanations with conceptual framing actually help you understand new domains. The lack of effectiveness validation means you're beta-testing whether concreteness metrics actually correlate with usefulness for your specific use cases. Also skip if your agents already respect 'be concise' system prompts—some developers report newer Claude models handle brevity instructions well enough that post-hoc filtering is unnecessary overhead.