PostHog's Architecture: How HogQL and Edge Evaluation Power a Full-Stack Analytics Monolith
Hook
PostHog stores session replays at 1/100th the size of video solutions by recording DOM mutations instead of pixels—then lets you query those sessions with SQL alongside your analytics events.
Context
The modern product stack has become absurdly fragmented. A typical SaaS company pays for Mixpanel (analytics), LaunchDarkly (feature flags), LogRocket (session replay), Sentry (error tracking), and Typeform (surveys)—then burns engineering time building data pipelines to connect them. Each tool has its own user identity system, pricing model, and query language. When your PM asks "what do users who hit this error actually do in their sessions?", you're manually correlating timestamps across three different platforms.
PostHog attempts to solve vendor sprawl through consolidation, not integration. Built on a Django/Python backend with React frontend, it's a genuine open-source monolith (MIT license for core features) that routes all product data—events, flags, replays, errors—through a unified ingestion pipeline into ClickHouse. The premise: if everything shares the same user IDs and timestamps in one analytical database, you can answer cross-domain questions with SQL instead of spreadsheet archaeology. Their recent push toward "self-driving products" adds LLM-powered diagnostics, but the real technical innovation is in how they've made a monolith scale without forcing you onto their cloud.
Technical Insight
PostHog's architecture splits into three layers: ingestion, storage, and query translation. Events enter through their SDKs, hit a Kafka queue, and flow through Node.js plugin-server workers that can run user-defined JavaScript transformations before writing to ClickHouse. This design choice—letting users inject code into the ingestion path—turns PostHog into a programmable analytics platform. You can enrich events with external API calls, filter spam, or route data to other systems without maintaining separate ETL infrastructure.
The query layer is where HogQL becomes critical. ClickHouse is phenomenally fast but notoriously difficult—its SQL dialect differs from PostgreSQL, property access requires understanding materialized columns, and joining event data with user properties involves arcane syntax. HogQL abstracts this complexity with a SQL-like language that compiles to optimized ClickHouse queries:
-- HogQL query (what you write)
SELECT
properties.$current_url as url,
person.properties.email,
count() as pageviews
FROM events
WHERE event = '$pageview'
AND timestamp > now() - interval 7 day
GROUP BY url, person.properties.email
HAVING pageviews > 10
This gets transpiled into a ClickHouse query that handles property extraction from JSON columns, joins the events table with the persons table on distinct_id, applies the correct sampling and partitioning, and uses materialized columns when available. The compiler understands PostHog's schema conventions—that properties.$current_url lives in a JSON column, that person.properties.email requires a join, that count() should leverage ClickHouse's aggregation functions.
The feature flag system takes a different approach to scale: aggressive client-side evaluation. Instead of hitting an API endpoint on every flag check (which would create millions of requests), SDKs fetch flag definitions with targeting rules and evaluate them locally:
# PostHog Python SDK - local evaluation
from posthog import Posthog
posthog = Posthog(
'your_api_key',
personal_api_key='your_personal_key', # Enables local evaluation
enable_local_evaluation=True
)
# This checks against locally cached flag rules—no API call
if posthog.feature_enabled('new-checkout-flow', 'user_123'):
show_new_checkout()
When you enable local evaluation, the SDK downloads all flag definitions and their targeting criteria (property filters, rollout percentages, cohorts) on initialization and polls for updates every 30 seconds. Flag checks become a local function call that evaluates boolean logic against user properties you provide. This reduces infrastructure costs dramatically—a high-traffic app making millions of flag checks generates only a few hundred API requests per hour for definition updates. The trade-off is staleness: flags can take up to 30 seconds to propagate, making this unsuitable for real-time experiments that need instant rollout changes.
Session replay demonstrates PostHog's pragmatic engineering. Rather than recording video (expensive storage, privacy nightmares, massive bandwidth), they use rrweb to capture DOM mutations as JSON events. When a user clicks a button, rrweb serializes what changed—button state, new elements rendered, CSS modifications—not pixels:
// What gets captured (simplified)
{
type: 3, // Mutation event
data: {
source: 2, // User interaction
type: 2, // MouseInteraction
id: 147, // DOM node ID
x: 523, y: 187,
adds: [{id: 148, node: {type: 'button', attributes: {class: 'primary'}}}],
removes: [{id: 142}]
}
}
These events compress beautifully (often 100KB for a 5-minute session) and get batched to S3/GCS with gzip. Playback reconstructs the DOM mutations client-side. The result: 10-100x smaller storage than video-based tools, with the ability to query replay metadata alongside analytics events because everything lives in ClickHouse. You can write SQL like SELECT session_id FROM events WHERE event = '$exception' AND properties.message LIKE '%payment%' then instantly pull up those sessions.
Their data warehouse sync feature closes the loop on external data. PostHog can treat your Stripe customers table or Salesforce opportunities as materialized views in ClickHouse, syncing them on a schedule. This enables joins between product events and business data:
-- Join product usage with Stripe subscription data
SELECT
stripe_customers.subscription_tier,
count(DISTINCT person_id) as active_users,
avg(session_duration) as avg_duration
FROM events
JOIN stripe_customers ON person.properties.email = stripe_customers.email
WHERE event = '$pageview'
GROUP BY subscription_tier
No ETL scripts, no reverse-ETL tools—just SQL. The sync engine is relatively simple: poll external APIs, write to ClickHouse, handle incremental updates. It works remarkably well for analytics use cases where eventual consistency is acceptable.
Gotcha
PostHog's "open source" positioning misleads many teams. While the code is MIT-licensed, self-hosting at scale requires deep expertise in ClickHouse clustering, Kafka tuning, and distributed systems. Their documentation explicitly states that self-hosted deployments beyond 100k events/month are unsupported—not because they're hostile to self-hosting, but because the operational complexity becomes prohibitive. You'll need to provision Kafka clusters, configure ClickHouse sharding, tune retention policies, and manage object storage lifecycles. Most teams attempting serious self-hosting end up migrating to PostHog Cloud after weeks of operations toil. The "open source" value is more about audit-ability and avoiding vendor lock-in than actual self-hosting viability.
The monolith architecture creates resource contention you can't easily solve. Session replay storage can grow to terabytes, heavy HogQL queries can peg CPU, and real-time event ingestion competes for I/O—all on the same infrastructure. PostHog Cloud handles this with internal sharding and workload isolation you can't replicate in their open-source deployment scripts. If you run it yourself, you'll face scaling cliffs where adding more events degrades query performance, or replay storage fills disks, and the only solution is vertical scaling or migrating to their managed offering. Additionally, HogQL's abstraction is leaky—once you need ClickHouse-specific optimizations like sampling strategies or exotic join types, you're writing raw ClickHouse SQL anyway, losing the simplicity that made HogQL appealing.
Verdict
Use if: You're a startup or mid-sized SaaS company currently paying for 3+ separate tools (analytics, flags, replay), your team has SQL-literate PMs who need ad-hoc analysis beyond dashboards, you're willing to commit to PostHog Cloud (not self-hosting), and you value consolidation over best-in-class features. The unified data model genuinely accelerates debugging when you can query "show me sessions where users hit errors after seeing feature flag X"—that cross-tool correlation is PostHog's killer feature. Skip if: You need true air-gapped self-hosting at scale (the operational burden will crush your team), you're already invested in best-of-breed tools like LaunchDarkly and Sentry with mature integrations, you have compliance requirements preventing third-party event streaming, or you need enterprise-grade SLAs for individual features (dedicated tools are more robust). PostHog is a productivity multiplier for the consolidation use case, but a migration nightmare if it doesn't fit your operational model.