Utopia: Building Enterprise AI with Bitemporal Knowledge Graphs
Hook
Most vector databases can tell you what they know right now. Utopia can tell you what it knew at any point in history—and when it learned it was wrong.
Context
Enterprise AI has a governance problem. When a loan application gets rejected, when a clinical trial decision gets questioned, when a legal discovery process demands 'show me what the system knew in March 2023'—vector databases and RAG systems fail catastrophically. They're append-only optimistically, but they have no concept of knowledge changing over time, no audit trail for decisions, no way to replay historical state.
This matters intensely in regulated environments. Financial services need to reconstruct loan decisions years after the fact. Pharmaceutical companies must prove what clinical trial data informed each protocol change. Legal teams building case chronologies can't use systems where facts silently update without provenance. The cloud-hosted knowledge graph vendors charge enterprise rates but force data offsite. Open-source alternatives are either pure vector stores with no temporal semantics (pgvector, Qdrant) or research databases that don't integrate document extraction and LLM orchestration. Utopia emerged to fill this gap: a self-hosted, bitemporal knowledge graph with RAG capabilities, designed for environments where 'explain this decision' means replaying exact historical state, not retrieving similar documents.
Technical Insight
Utopia's architecture centers on bitemporal modeling—every fact records two timestamps. Transaction time captures when the system learned something ('we recorded this entity on 2024-01-15'). Valid time captures when something was true in the world ('this person held CFO title from 2023-06 to 2024-01'). This distinction is critical for auditable AI. When you query 'what did we know about XYZ Corp on March 1st?', the system reconstructs graph state using facts where transaction time ≤ 2024-03-01 and valid time overlaps that date.
The ingestion pipeline demonstrates how ontology-as-constraint works. Documents flow through extraction that generates entities conforming to loaded ontologies (schema.org, FOAF, custom vocabularies). Here's a simplified extraction workflow:
// Entity extraction with ontology constraints
let extraction = ExtractorPipeline::new()
.with_ontology("schema.org")
.with_llm_endpoint("http://localhost:11434"); // Ollama local
let doc = Document::from_file("contracts/acme_2024.pdf")?;
let entities = extraction.extract(doc).await?;
for entity in entities {
// Each entity has types from the ontology
match entity.entity_type {
OntologyType::Organization => {
// Axioms define constraints
if entity.relations.get("parentOrganization").is_some() {
// Transitivity axiom auto-derives ancestor relations
inference_engine.derive_transitive(
&entity,
"parentOrganization",
"ancestorOrganization"
);
}
},
OntologyType::Person => {
// Cardinality constraints from ontology
let employer_count = entity.relations
.get("worksFor")
.map(|r| r.len())
.unwrap_or(0);
if employer_count > 3 {
// Violates schema.org cardinality expectation
review_queue.push(ReviewTask {
entity: entity.id,
issue: CardinalityViolation {
property: "worksFor",
expected_max: 3,
actual: employer_count,
},
valid_time: Instant::now(),
transaction_time: SystemTime::now(),
});
}
},
_ => {}
}
}
Notice the forward-chaining inference deriving ancestor relationships from parent relationships. The ontology defines transitivity axioms that compile to these derivation rules. When conflicts arise—say, a symmetry violation where A worksFor B but B doesn't workFor A when a 'colleague' relation should be symmetric—the system doesn't guess. It creates a human review task.
The three-stage entity resolution handles the hard problem of determining whether 'Acme Corp' in document A is the same as 'ACME Corporation' in document B. Stage one: exact string match on canonical names. Stage two: embedding similarity above a threshold (0.85 default) using pgvector's cosine distance. Stage three: LLM judgment with full context. Critically, every merge decision is recorded in the decision ledger with full entity snapshots, making it reversible:
// Entity resolution with full undo capability
let resolution = EntityResolver::new()
.with_embedding_threshold(0.85)
.with_llm_judge(llm_endpoint);
let candidate_matches = resolution
.find_candidates(new_entity, &graph)
.await?;
for candidate in candidate_matches {
let decision = resolution.judge_match(
&new_entity,
&candidate,
JudgmentContext {
surrounding_facts: graph.neighbors(&candidate, 2),
document_sources: vec![doc.id],
}
).await?;
if decision.confidence > 0.9 {
// Record merge in decision ledger before executing
decision_ledger.record(Decision {
decision_type: EntityMerge,
source_snapshot: new_entity.clone(),
target_snapshot: candidate.clone(),
rationale: decision.reasoning,
actor: Actor::System("entity_resolver_v2"),
transaction_time: SystemTime::now(),
});
graph.merge_entities(new_entity.id, candidate.id)?;
}
}
The decision ledger outlives the entities. Even if you delete an entity from the graph, the ledger preserves every merge, split, or confirmation decision with full object state. This is the foundation for temporal queries like 'reconstruct the entity graph as it existed when we made loan decision #4821'.
Ontology2SQL is where things get genuinely interesting. Traditional text-to-SQL systems struggle with semantic ambiguity—'revenue' could be a column in the financials table or the sales table or need to be computed from line items. Utopia first maps database schemas onto the knowledge graph's ontology, then generates SQL respecting those semantic relationships:
-- Database schema gets mapped to ontology types
CREATE TABLE companies (
id UUID PRIMARY KEY,
name TEXT,
-- Mapped to schema.org/Organization
);
CREATE TABLE financial_reports (
id UUID PRIMARY KEY,
company_id UUID REFERENCES companies(id),
total_revenue DECIMAL,
report_year INT,
-- Mapped to schema.org/FinancialProduct
);
When a user asks 'What was Acme's revenue in 2023?', the system:
- Resolves 'Acme' to entities in the knowledge graph (handles 'Acme Corp' vs 'ACME Corporation')
- Maps 'revenue' to ontology property schema:totalRevenue
- Finds database columns mapped to that property (financial_reports.total_revenue)
- Generates SQL joining through ontology-defined relationships:
SELECT fr.total_revenue, fr.report_year
FROM companies c
JOIN financial_reports fr ON c.id = fr.company_id
WHERE c.name ILIKE '%acme%'
AND fr.report_year = 2023;
The power is that this query can span both structured data (the financial_reports table) and extracted documents (contract PDFs mentioning revenue figures) because they share ontology types. The graph knows that revenue mentions in contracts and revenue columns in databases both map to schema:totalRevenue.
The deployment model is deliberately minimal. Tantivy (a Rust full-text search library) runs in-process—no separate Elasticsearch cluster. pgvector handles embeddings in Postgres. Job queues are Postgres tables with SELECT FOR UPDATE SKIP LOCKED for concurrency. The entire stack runs air-gapped with local LLM endpoints like Ollama. This isn't academic—it's the only architecture that works when data legally cannot leave infrastructure or touch third-party APIs.
Gotcha
The timestamp precision ceiling is day-level granularity. If you're ingesting high-frequency event data where order within a day matters, Utopia will lose causality. The system rounds UTC timestamps to midnight, which can shift events across date boundaries—a transaction at 11:59 PM and another at 12:01 AM might appear in wrong order after rounding.
Forward-chaining inference is off by default because one bad axiom poisons the entire derived fact set. The documentation warns about this, but there's no tooling for axiom validation or sandboxed inference testing. You can't say 'test these transitivity rules on a sample subgraph before unleashing them'. Either you enable inference globally and risk deriving thousands of wrong facts, or you skip it entirely and lose the semantic reasoning capabilities that make ontologies valuable. The middle ground—careful axiom testing—requires manual SQL queries and graph exports.
The single Rust binary architecture means vertical scaling only. You cannot horizontally partition extraction jobs across nodes. One document queue, one embedding generator, one query executor. For continuous high-volume ingestion (thousands of documents per hour), you'll hit throughput limits. The authors acknowledge this is v0.1 software with evolving schemas—database migrations only roll forward with no rollback support, meaning production upgrades require full backups and acceptance that downgrades are impossible.
Verdict
Use if: You're building agent systems in regulated industries where 'explain this decision' means replaying exact historical state, not just retrieving similar documents. If you need auditable AI for financial services, pharmaceuticals, legal tech, or government contracts where data cannot leave infrastructure and every knowledge change needs provenance. If you're comfortable running v0.x software and contributing fixes upstream, because the bitemporal graph with decision ledger is genuinely novel infrastructure that doesn't exist elsewhere in open source. Skip if: You need real-time streaming ingestion at scale (day-precision timestamps and single-binary architecture won't support high-frequency data), horizontal scaling beyond one powerful Postgres instance, or production-hardened stability with schema guarantees. If you're just building better search over documents, pgvector + LlamaIndex is simpler. Utopia is for teams where knowledge graphs aren't nice-to-have—they're compliance requirements.