> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Phoenix: Building LLM Observability on OpenTelemetry's Foundation

[ View on GitHub ]

Phoenix: Building LLM Observability on OpenTelemetry's Foundation

Hook

Your AI coding assistant can now query your production LLM traces directly through natural language—Phoenix just turned observability dashboards into agent-native infrastructure, and it's built entirely on open standards.

Context

Debugging LLM applications is fundamentally different from debugging traditional software. When your RAG pipeline returns irrelevant documents or your agent halts mid-task, there's no stack trace pointing to line 47. The failure exists in the semantic space—bad retrieval, prompt drift, context window overflow, or the model simply having a bad day. Early LLM builders resorted to print statements and prayer, dumping prompts to log files and manually correlating outputs with inputs across dozens of API calls.

The first wave of LLM observability tools were vendor-specific dashboards tied to specific frameworks. LangSmith locked you into LangChain's ecosystem. Helicone proxied OpenAI calls but couldn't see inside your retrieval logic. What was missing was a universal instrumentation layer that could trace across frameworks, store data locally, and avoid vendor lock-in. Phoenix emerged from Arize AI's ML monitoring team with a bet on OpenTelemetry as that universal layer—the same standard that traces microservices could trace multi-step LLM workflows, with Phoenix providing LLM-specific indexing, evaluation, and visualization on top.

Technical Insight

Phoenix's architecture centers on OpenInference, a collection of auto-instrumentors that hook into framework internals at import time and emit OpenTelemetry spans. When you instrument a LangChain application, OpenInference detects chain invocations and automatically creates hierarchical spans with framework-aware attributes—no manual span creation required:

from openinference.instrumentation.langchain import LangChainInstrumentor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor

# Configure OpenTelemetry to send spans to Phoenix
tracer_provider = TracerProvider()
tracer_provider.add_span_processor(
    SimpleSpanProcessor(OTLPSpanExporter("http://localhost:6006/v1/traces"))
)

# One-line instrumentation for all LangChain calls
LangChainInstrumentor().instrument(tracer_provider=tracer_provider)

# Your existing LangChain code remains unchanged
from langchain.chains import RetrievalQA
from langchain.vectorstores import Chroma

qa_chain = RetrievalQA.from_chain_type(
    llm=ChatOpenAI(model="gpt-4"),
    retriever=Chroma.from_documents(docs).as_retriever()
)

# This call automatically generates spans for retrieval, LLM invocation, and chain execution
result = qa_chain.run("What are the main findings?")

The instrumentation captures not just timing data but semantic attributes specific to LLM operations. Each span includes token counts, model names, input/output messages, and for retrieval spans, the actual documents retrieved with their similarity scores. Phoenix indexes these attributes in its storage layer, enabling queries like "show me all RAG calls where we retrieved more than 10 chunks but the final answer had low relevance."

The evaluation engine builds on this trace data with LLM-as-judge workflows. Phoenix ships with templates for hallucination detection, relevance scoring, and RAG quality checks that run as batch jobs against stored traces. The clever part is the two-tier approach: smaller models (GPT-3.5-turbo) handle initial scoring for cost efficiency, while GPT-4-class models generate explanations only when needed. The prompt engineering includes reference answers and rubrics to constrain the judge's hallucination:

from phoenix.evals import (
    HallucinationEvaluator,
    RAG_RELEVANCY_PROMPT_TEMPLATE,
    run_evals
)
import pandas as pd

# Load traces from Phoenix as a dataframe
traces_df = px.Client().get_spans_dataframe(
    filter_condition="span_kind == 'RETRIEVER'"
)

# Run hallucination detection across all retrieval spans
hallucination_eval = HallucinationEvaluator(
    model="gpt-4-turbo-preview"
)

eval_results = run_evals(
    dataframe=traces_df,
    evaluators=[hallucination_eval],
    provide_explanation=True  # Generates reasoning for each score
)

# Results include numeric scores and natural language explanations
print(eval_results[["score", "explanation"]].head())

The MCP (Model Context Protocol) integration represents a paradigm shift in how developers interact with observability data. Instead of switching to a web dashboard, you can query traces directly from Claude or Cursor:

# Start Phoenix with MCP server enabled
import phoenix as px

px.launch_app(enable_mcp=True)

# In your AI coding assistant (Claude Desktop, Cursor):
# "Show me all LangChain traces from the last hour where latency exceeded 5 seconds"
# "What were the retrieved documents for trace ID abc123?"
# "Compare token usage between gpt-4 and gpt-3.5-turbo calls today"

The MCP server exposes Phoenix's GraphQL API through a conversational interface, letting agents construct queries based on natural language requests. This closes the loop from observation to action—your coding assistant can identify a problematic trace, retrieve the exact prompt and context, and suggest fixes without you leaving the editor.

Phoenix's dataset management treats prompt variations as first-class versioned entities. You can snapshot a set of test inputs, run evaluations across multiple prompt templates, and compare results side-by-side. This brings A/B testing discipline to prompt engineering, replacing gut-feel iterations with systematic measurement. The dataset API lets you pull examples from production traces, curate them into test sets, and lock them as immutable versions for regression testing as your prompts evolve.

Gotcha

The auto-instrumentation approach creates brittle dependencies on framework internals. LangChain refactors its chain abstractions regularly, and OpenInference instrumentors lag behind by weeks. During those gaps, you're either pinning to outdated framework versions or manually instrumenting spans—exactly what auto-instrumentation was supposed to avoid. The LlamaIndex instrumentor broke entirely during the 0.9 to 0.10 migration and took a month to catch up.

Storage scaling is Phoenix's Achilles heel. SQLite works fine for local development, but production deployments need PostgreSQL, and the documentation provides minimal guidance on tuning for high-throughput trace ingestion. There's no discussion of partitioning strategies, retention policies, or how to handle thousands of spans per minute without table bloat. Teams hitting scale issues report query performance degrading after a few million spans, requiring manual database maintenance that shouldn't be necessary for a 2024 observability platform. The evaluation framework is Python-only, forcing TypeScript applications to run evals in a separate service or skip LLM-based quality checks entirely—a major gap given the Node.js LLM ecosystem's growth.

Verdict

Use Phoenix if you're building production LLM applications with Python frameworks (LangChain, LlamaIndex, OpenAI SDK), need observability without vendor lock-in, and have ops capacity to run stateful services. The OpenTelemetry foundation provides a real migration path to commercial backends later, and the evaluation templates genuinely accelerate RAG quality work. The MCP integration is legitimately innovative for teams using AI coding assistants. Skip Phoenix if you're pure TypeScript (evaluation tooling won't work), lack infrastructure to run PostgreSQL for production traces, or need real-time trace queries for user-facing features—it's built for post-hoc analysis. Also skip if you're in regulated environments concerned about AI assistants querying production data; the MCP authentication story needs hardening. For simple OpenAI API monitoring, Helicone's proxy approach is simpler. For teams wanting zero ops burden, LangSmith's hosted offering makes sense despite vendor lock-in. But for data ownership, multi-framework support, and open-source insurance, Phoenix is the strongest play.