The decision desk
LLM observability tools: Langfuse vs LangSmith vs Arize Phoenix
Compare Langfuse, LangSmith, and Arize Phoenix for inspecting AI application traces, connecting prompts to results, and choosing a hosted or self-hosted workflow.
Our take
Choose Langfuse for integrated tracing and prompt workflows, LangSmith for a managed agent-development platform, and Phoenix for a self-hostable analysis environment.
The quick difference
3 tools, at a glance
Start with the fit. Read the individual breakdowns below before choosing. Scroll the table on smaller screens.
| Tool | Best fit | Main strengths | What to check |
|---|---|---|---|
| Langfuse | Engineering teams combining AI observability and prompt iteration across frameworks. | Model, retrieval, and tool-call traces · Prompt versioning linked to application runs · Evaluations, datasets, and experiment comparisons | Hosted usage is billed in Langfuse units, not a flat price per end-user conversation. |
| LangSmith | Teams wanting managed tracing alongside LangChain ecosystem development and deployment. | Detailed traces and conversation views · Dashboards, alerts, and feedback annotation · Evaluation and prompt-engineering workflows | Seat subscriptions and usage charges are separate. |
| Arize Phoenix | Engineers wanting a locally runnable tracing and evaluation environment. | OpenTelemetry and OpenInference tracing · Prompt playground and span replay · Code, model-based, and human evaluations | Self-hosting requires infrastructure and operations even without a software fee. |
Compare pricing, features, and source dates
Prices describe the named plan, not total cost. Different billing units are not directly equivalent. Scroll the table horizontally on smaller screens.
| At a glance | LangfuseVendor documentation | LangSmithVendor documentation | Arize PhoenixVendor documentation |
|---|---|---|---|
| Consider it for | Engineering teams combining AI observability and prompt iteration across frameworks. | Teams wanting managed tracing alongside LangChain ecosystem development and deployment. | Engineers wanting a locally runnable tracing and evaluation environment. |
| Named plan / entry point | Core: $29/month plus excess usage USD per month plus metered units | Plus: $39/seat/month plus usage USD per seat per month plus metered usage | Free to self-host; infrastructure extra Self-hosted software; hosting and model costs separate |
| Free option | Available | Available | Available |
| Documented features |
|
|
|
| Limitations |
|
|
|
| Integrations | Python SDK, JavaScript SDK, OpenTelemetry, LangChain, LlamaIndex, LiteLLM | LangChain, OpenAI, Anthropic, CrewAI, Vercel AI SDK, Pydantic AI | OpenTelemetry, OpenInference, LangChain, LlamaIndex, OpenAI, Anthropic |
| Billing details | Hobby includes 50,000 units/month, two users, and 30-day data access. Core includes 100,000 units/month, unlimited users, and 90-day access; listed additional usage starts at $8 per 100,000 units. Model-provider charges are separate. Self-hosting is also available. | Developer has one free seat and up to 5,000 base traces per month, then pay-as-you-go. Plus includes up to 10,000 base traces per month, then usage charges. Deployment and other services have separate metering; trace allowances are not model-token allowances. | Official documentation states that Phoenix is free to self-host without feature gates or software usage limits. This does not include infrastructure or external evaluation-model fees. Managed Arize AX and its commercial services are separate. |
| Last checked | Sep 19, 2026 Official source | Sep 19, 2026 Official source | Sep 19, 2026 Official source |
Seeing what happened inside an AI application
An AI response may involve a model call, document retrieval, and several tool calls. Langfuse, LangSmith, and Arize Phoenix make those steps visible and connect recorded behavior to ways of improving it. This comparison focuses on their tracing experience, iteration tools, and hosting models rather than claiming a benchmark winner.
Langfuse
Langfuse brings tracing, prompt management, and evaluations into one AI engineering platform. A trace can include model calls and the surrounding retrieval or application logic, with cost and latency information helping explain how a run behaved. Sessions connect multi-turn conversations, while prompt versions can be linked to the traces they produced. That relationship is useful when a team wants to understand whether a changed instruction improved responses or simply increased expense. The platform also provides datasets, experiments, and several ways to attach quality scores. Native SDKs, framework integrations, and OpenTelemetry give it multiple entry points into an existing stack. Langfuse offers both hosted plans and self-hosting; the cloud tiers differ in usage allowances, history access, and administrative features.
LangSmith
LangSmith provides detailed application traces alongside monitoring, feedback, evaluation, and broader agent-development services. Its observability documentation describes conversation and run views, dashboards, alerts, and workflows that turn recorded behavior into evaluation datasets. Although it comes from the LangChain ecosystem, its integrations also cover other frameworks and providers, so it should not be understood as requiring a LangChain-only application. The attraction is a managed environment where a team can investigate a production run and continue into related development workflows. For teams already considering LangChain’s wider platform, that continuity can simplify the product choice. Pricing separates seat access from usage, and deployment or other services introduce additional metering. Enterprise self-hosted and hybrid options are commercial offerings rather than a free community edition.
Arize Phoenix
Arize Phoenix is a self-hostable workspace for tracing and improving AI applications. It accepts OpenTelemetry traces and uses OpenInference instrumentation to describe model calls, retrieval, tool use, and custom logic. Engineers can inspect those steps, replay model calls with changed inputs, and work with prompt versions and datasets. Its evaluation tools support model-based checks, code checks, and human labels, while experiments compare application changes on shared inputs. Phoenix is particularly attractive when the team wants a locally runnable environment that can grow into its own deployment. The documentation states that self-hosting has no software fee or feature gates, although infrastructure and evaluation-model usage still cost money. Arize AX is a separate managed enterprise platform and should not be treated as the same purchase.
Tracing and the connection to prompt changes
All three can show more than the final model response, so the useful distinction is what the team does after opening a trace. Langfuse explicitly connects prompt versions, sessions, and performance metrics in its platform. LangSmith connects production traces with feedback, evaluation datasets, and its wider agent-development services. Phoenix emphasizes an inspect-and-iterate environment with span replay, a prompt playground, and experiments. For a team investigating a changed instruction, Langfuse’s prompt-to-trace relationship is appealing. For a team working across a managed development stack, LangSmith offers continuity. For engineers who want to replay a specific model call and explore changes in their own environment, Phoenix is a strong candidate.
Framework support and where the platform runs
Langfuse and Phoenix both provide self-hostable paths, while LangSmith offers commercial Enterprise self-hosted and hybrid deployments. Langfuse and Phoenix also document OpenTelemetry-based ingestion, making them relevant when an application combines several frameworks. LangSmith supports a range of providers and frameworks beyond LangChain through its own documented integrations. Our choice would depend on whether the team wants to operate the observability service or buy a managed experience. A technical team seeking a freely self-hosted analysis environment has a direct reason to consider Phoenix. A team wanting cloud collaboration can compare Langfuse and LangSmith around their included users, usage model, and surrounding workflow.
Comparing seats, units, and hosting costs
Langfuse’s cloud plans use metered units and defined history-access periods. LangSmith separates seats from trace and other service usage. Phoenix’s free self-hosted software shifts the bill toward infrastructure and any external models used for evaluation. Those numbers are not interchangeable: a conversation can contain multiple recorded operations, and an evaluation can generate model charges outside the observability subscription.
Which LLM observability tool should you choose?
Choose Langfuse for a connected tracing and prompt workflow, LangSmith for a managed platform alongside agent development, or Phoenix for a self-hostable investigation environment. Each supports meaningful iteration; the best fit depends on the team’s stack and preferred operating model.
Sources & methodology
This article uses vendor documentation and our editorial analysis. We have not performed a controlled hands-on test. Product facts were checked on Sep 19, 2026; each profile records its own verification date.
- Langfuse profile and source notes · Platform documentation · Official pricing · Observability documentation
- LangSmith profile and source notes · Observability documentation · Official pricing · Usage and billing
- Arize Phoenix profile and source notes · Product overview · Platform documentation · Self-hosting and cost