> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

OpenHuman: Why the Fastest Rust Agent Harness Ships with Vendor Lock-In

[ View on GitHub ]

OpenHuman: Why the Fastest Rust Agent Harness Ships with Vendor Lock-In

Hook

OpenHuman claims to host agents at 1.77 MiB each versus 48 MiB for separate processes—a 25x density improvement that would transform edge deployment economics. But the 'local-first' marketing hides a dependency on TinyHumans cloud services for core decision-making.

Context

Agent orchestration frameworks have converged on a bloated pattern: Python runtimes with LangChain or LangGraph abstractions, separate process-per-agent isolation, and cloud-native memory stores that assume unlimited resources. AutoGPT spawns Docker containers. CrewAI layers role abstractions atop LangChain's already-heavy SDK. Semantic Kernel brings .NET's ceremony to what should be lightweight coordination logic. Each approach optimizes for developer ergonomics in a cloud environment where memory and latency don't constrain architecture.

This works poorly for three emerging deployment targets: desktop applications where users expect sub-second response and local execution, edge devices with tight memory budgets, and SaaS platforms hosting thousands of isolated agent instances. OpenHuman enters as a Rust-native alternative that embeds agents in-process, uses feature gates to compile only needed capabilities, and serializes memory as portable markdown files. The TinyHumans team built it after frustration with the resource overhead of running multiple Python-based agents in their own products—a classic 'scratch your own itch' origin story that produced genuinely different architectural choices.

Technical Insight

Memory Backends

Loadable Modules

OpenHuman Core Runtime

Frontend Layer

Multi-Agent Space

Memory Trees

Memory Trees

Tauri Desktop App

Ratatui TUI

Web SPA

openhuman-embed API

Shared Rust Runtime

Agent 1

Provider + Memory + Sandbox

Agent 2

Provider + Memory + Sandbox

Capability Bus

*-bus contracts

Jev Decision Proxy

System One API

TinyFlows Engine

22 node types

tinydocs

tinyvoice

tinyjuice

tinyruntime

Local Disk

Obsidian Vault

TinyCortex Remote

System architecture — auto-generated

OpenHuman's core bet is that separate processes per agent are unnecessary overhead. Instead, a single Rust runtime hosts multiple independent agents sharing the same address space but with isolated providers, memory backends, and sandboxed working directories. The openhuman-embed library exposes this through a clean API:

use openhuman_embed::{Runtime, AgentConfig, MemoryBackend};

let mut runtime = Runtime::new()?;

let agent_a = runtime.spawn_agent(AgentConfig {
    name: "researcher",
    provider: "openai/gpt-4o",
    memory: MemoryBackend::Local("./memory/researcher"),
    skills: vec!["web_search", "document_reader"],
    sandbox: true,
})?;

let agent_b = runtime.spawn_agent(AgentConfig {
    name: "writer",
    provider: "anthropic/claude-3-5-sonnet",
    memory: MemoryBackend::TinyCortex { remote: true },
    skills: vec!["markdown_generator"],
    sandbox: true,
})?;

// Agents share the runtime but can't access each other's memory
let result = agent_a.turn("Research Rust async patterns").await?;
agent_b.turn(format!("Write a tutorial based on: {}", result)).await?;

The memory density claim comes from this architecture: each agent adds roughly 1.77 MiB at scale because they share loaded libraries, the tokio runtime, and compiled skill modules. Separate processes would duplicate all of this, plus incur OS scheduler overhead and inter-process communication costs. The tradeoff is reduced fault isolation—an agent that panics can theoretically crash the entire runtime, though the sandboxing and Rust's memory safety mitigate this significantly.

Cargo feature gates slice the monolith into optional capabilities. The repository defines nine contributor gates (the full feature matrix for development) and a product-features.txt manifest that end users compose:

# Minimal build: just chat with an LLM
[dependencies]
openhuman-embed = { version = "0.1", default-features = false }

# Add web search and document skills
openhuman-embed = { version = "0.1", features = ["skills-web", "skills-docs"] }

# Full desktop app capabilities
openhuman-embed = { version = "0.1", features = ["full"] }

This produces binary size differences that matter for distribution: 51 MiB stripped for a pure chat agent, 116 MiB unstripped with workflow engine, media processing, MCP server support, and HTTP endpoints. For comparison, bundling a Python agent harness with embedded interpreter and dependencies rarely gets below 200 MiB. This is library-grade modularity typically reserved for embedded systems or kernel components, not desktop applications.

The Jev subsystem is where architectural pragmatism meets vendor dependency. Instead of asking an LLM to generate JSON tool calls—which costs tokens, adds latency, and requires parsing unreliable free-text output—OpenHuman routes structured decisions through a specialized ranking model. When an agent needs to select from 215 built-in tools plus 1,000 Composio API actions, Jev scores candidates and returns the top match with 62% top-1 accuracy. This is dramatically faster (sub-10ms versus 500ms+ for an LLM generation) and cheaper (pennies per thousand decisions versus dollars). The catch: Jev runs exclusively through the TinyHumans System One API. Without credentials, OpenHuman falls back to BM25 keyword search at 22.5% accuracy—a 64% degradation that makes tool-heavy workflows nearly unusable.

Memory Trees serialize as Obsidian vaults—plain markdown files with YAML frontmatter and wikilink references. An agent's memory about a user's project preferences might live at memory/researcher/projects/rust-learning.md with content like:

---
type: preference
created: 2025-01-15T10:23:11Z
tags: [rust, learning, async]
---

# Rust Learning Preferences

User prefers code-heavy explanations over theory.
Interested in async patterns, particularly tokio and async-std comparisons.
Finds trait bounds confusing—needs concrete examples.

## Related
- [[async-runtime-comparison]]
- [[trait-bound-examples]]

This dual representation—structured graph internally, portable files externally—solves the memory portability problem that plagues cloud platforms. If TinyHumans disappears tomorrow, users keep readable markdown files that any other tool can parse. The workflow canvas extends this pattern to automation: agents draft tinyflows graphs (a separate open-source workflow runtime with 22 node types) that users review and approve rather than wiring nodes manually. It inverts the n8n/Zapier model by making intent description the interface and node graphs the output artifact.

Gotcha

The 'local-first' and 'open-source' branding misleads about operational reality. Jev decision routing—critical for multi-tool agents—requires TinyHumans API credentials with no local inference option. The company provides 'one API key for everything' (LLM routing, embeddings, web search, voice synthesis, Jev scoring) which simplifies integration but creates total dependency on their infrastructure. The Apache 2.0 license means you can fork and modify the Rust code, but without reimplementing Jev or accepting degraded tool selection, you're operationally dependent on a venture-funded startup's continued service availability and pricing.

The performance numbers lack reproducibility details that matter for evaluation. Bootstrap time of 476ms and cold turn latency of 102ms are impressive, but the benchmarks don't specify hardware (M3 Max versus AWS t3.medium produce very different results), model size (GPT-4o versus GPT-4o-mini have different processing overhead), context length (1K versus 32K tokens dramatically affect latency), or concurrent agent count under load. The 40,412 stars accumulated in roughly a week suggest viral marketing success, not battle-tested production validation. The repository warns 'early beta' which is accurate—this is a promising v0.1 with excellent architectural foundations but minimal production deployments beyond the reference desktop application.

Verdict

Use if: You're embedding agent orchestration in a Rust product where memory density and cold-start latency justify tight integration, you're building desktop or edge applications where users expect local execution with cloud augmentation rather than pure cloud dependence, you want opinionated memory and workflow patterns (Memory Trees, tinyflows canvas) rather than building these abstractions yourself, and you're comfortable with TinyHumans vendor dependency for decision routing and managed services. The openhuman-embed library is legitimately excellent for these constraints. Skip if: You need ecosystem maturity and breadth—LangChain has orders of magnitude more integrations despite OpenHuman's marketing claims about tool coverage, you require vendor independence for regulated industries or long-term operational control, you're evaluating based on production scale evidence rather than synthetic benchmarks from a one-week-old project, or your team doesn't have Rust expertise to debug issues in a pre-1.0 harness with limited community knowledge base. The architectural ideas are sound but the project needs six months of production hardening before it's suitable for critical workloads.