The decision desk

LLM observability tools: Langfuse vs LangSmith vs Arize Phoenix

Compare Langfuse, LangSmith, and Arize Phoenix for inspecting AI application traces, connecting prompts to results, and choosing a hosted or self-hosted workflow.

By AIPicksy EditorialUpdated Sep 19, 20264 min readSource-based editorial

Our take

Choose Langfuse for integrated tracing and prompt workflows, LangSmith for a managed agent-development platform, and Phoenix for a self-hostable analysis environment.

The quick difference

3 tools, at a glance

Start with the fit. Read the individual breakdowns below before choosing. Scroll the table on smaller screens.

Product strengths and trade-offs
ToolBest fitMain strengthsWhat to check
LangfuseEngineering teams combining AI observability and prompt iteration across frameworks.Model, retrieval, and tool-call traces · Prompt versioning linked to application runs · Evaluations, datasets, and experiment comparisonsHosted usage is billed in Langfuse units, not a flat price per end-user conversation.
LangSmithTeams wanting managed tracing alongside LangChain ecosystem development and deployment.Detailed traces and conversation views · Dashboards, alerts, and feedback annotation · Evaluation and prompt-engineering workflowsSeat subscriptions and usage charges are separate.
Arize PhoenixEngineers wanting a locally runnable tracing and evaluation environment.OpenTelemetry and OpenInference tracing · Prompt playground and span replay · Code, model-based, and human evaluationsSelf-hosting requires infrastructure and operations even without a software fee.
Compare pricing, features, and source dates

Prices describe the named plan, not total cost. Different billing units are not directly equivalent. Scroll the table horizontally on smaller screens.

Side-by-side facts from the current product profiles
At a glanceLangfuseVendor documentationLangSmithVendor documentationArize PhoenixVendor documentation
Consider it forEngineering teams combining AI observability and prompt iteration across frameworks.Teams wanting managed tracing alongside LangChain ecosystem development and deployment.Engineers wanting a locally runnable tracing and evaluation environment.
Named plan / entry pointCore: $29/month plus excess usage
USD per month plus metered units
Plus: $39/seat/month plus usage
USD per seat per month plus metered usage
Free to self-host; infrastructure extra
Self-hosted software; hosting and model costs separate
Free optionAvailableAvailableAvailable
Documented features
  • Model, retrieval, and tool-call traces
  • Prompt versioning linked to application runs
  • Evaluations, datasets, and experiment comparisons
  • Detailed traces and conversation views
  • Dashboards, alerts, and feedback annotation
  • Evaluation and prompt-engineering workflows
  • OpenTelemetry and OpenInference tracing
  • Prompt playground and span replay
  • Code, model-based, and human evaluations
Limitations
  • Hosted usage is billed in Langfuse units, not a flat price per end-user conversation.
  • Data access periods and enterprise controls vary by tier.
  • Seat subscriptions and usage charges are separate.
  • Self-hosted and hybrid offerings are commercial Enterprise options.
  • Self-hosting requires infrastructure and operations even without a software fee.
  • Arize AX is a separate managed enterprise product; its pricing is not Phoenix pricing.
IntegrationsPython SDK, JavaScript SDK, OpenTelemetry, LangChain, LlamaIndex, LiteLLMLangChain, OpenAI, Anthropic, CrewAI, Vercel AI SDK, Pydantic AIOpenTelemetry, OpenInference, LangChain, LlamaIndex, OpenAI, Anthropic
Billing detailsHobby includes 50,000 units/month, two users, and 30-day data access. Core includes 100,000 units/month, unlimited users, and 90-day access; listed additional usage starts at $8 per 100,000 units. Model-provider charges are separate. Self-hosting is also available.Developer has one free seat and up to 5,000 base traces per month, then pay-as-you-go. Plus includes up to 10,000 base traces per month, then usage charges. Deployment and other services have separate metering; trace allowances are not model-token allowances.Official documentation states that Phoenix is free to self-host without feature gates or software usage limits. This does not include infrastructure or external evaluation-model fees. Managed Arize AX and its commercial services are separate.
Last checkedSep 19, 2026
Official source
Sep 19, 2026
Official source
Sep 19, 2026
Official source

Seeing what happened inside an AI application

An AI response may involve a model call, document retrieval, and several tool calls. Langfuse, LangSmith, and Arize Phoenix make those steps visible and connect recorded behavior to ways of improving it. This comparison focuses on their tracing experience, iteration tools, and hosting models rather than claiming a benchmark winner.

Langfuse

Langfuse brings tracing, prompt management, and evaluations into one AI engineering platform. A trace can include model calls and the surrounding retrieval or application logic, with cost and latency information helping explain how a run behaved. Sessions connect multi-turn conversations, while prompt versions can be linked to the traces they produced. That relationship is useful when a team wants to understand whether a changed instruction improved responses or simply increased expense. The platform also provides datasets, experiments, and several ways to attach quality scores. Native SDKs, framework integrations, and OpenTelemetry give it multiple entry points into an existing stack. Langfuse offers both hosted plans and self-hosting; the cloud tiers differ in usage allowances, history access, and administrative features.

LangSmith

LangSmith provides detailed application traces alongside monitoring, feedback, evaluation, and broader agent-development services. Its observability documentation describes conversation and run views, dashboards, alerts, and workflows that turn recorded behavior into evaluation datasets. Although it comes from the LangChain ecosystem, its integrations also cover other frameworks and providers, so it should not be understood as requiring a LangChain-only application. The attraction is a managed environment where a team can investigate a production run and continue into related development workflows. For teams already considering LangChain’s wider platform, that continuity can simplify the product choice. Pricing separates seat access from usage, and deployment or other services introduce additional metering. Enterprise self-hosted and hybrid options are commercial offerings rather than a free community edition.

Arize Phoenix

Arize Phoenix is a self-hostable workspace for tracing and improving AI applications. It accepts OpenTelemetry traces and uses OpenInference instrumentation to describe model calls, retrieval, tool use, and custom logic. Engineers can inspect those steps, replay model calls with changed inputs, and work with prompt versions and datasets. Its evaluation tools support model-based checks, code checks, and human labels, while experiments compare application changes on shared inputs. Phoenix is particularly attractive when the team wants a locally runnable environment that can grow into its own deployment. The documentation states that self-hosting has no software fee or feature gates, although infrastructure and evaluation-model usage still cost money. Arize AX is a separate managed enterprise platform and should not be treated as the same purchase.

Tracing and the connection to prompt changes

All three can show more than the final model response, so the useful distinction is what the team does after opening a trace. Langfuse explicitly connects prompt versions, sessions, and performance metrics in its platform. LangSmith connects production traces with feedback, evaluation datasets, and its wider agent-development services. Phoenix emphasizes an inspect-and-iterate environment with span replay, a prompt playground, and experiments. For a team investigating a changed instruction, Langfuse’s prompt-to-trace relationship is appealing. For a team working across a managed development stack, LangSmith offers continuity. For engineers who want to replay a specific model call and explore changes in their own environment, Phoenix is a strong candidate.

Framework support and where the platform runs

Langfuse and Phoenix both provide self-hostable paths, while LangSmith offers commercial Enterprise self-hosted and hybrid deployments. Langfuse and Phoenix also document OpenTelemetry-based ingestion, making them relevant when an application combines several frameworks. LangSmith supports a range of providers and frameworks beyond LangChain through its own documented integrations. Our choice would depend on whether the team wants to operate the observability service or buy a managed experience. A technical team seeking a freely self-hosted analysis environment has a direct reason to consider Phoenix. A team wanting cloud collaboration can compare Langfuse and LangSmith around their included users, usage model, and surrounding workflow.

Comparing seats, units, and hosting costs

Langfuse’s cloud plans use metered units and defined history-access periods. LangSmith separates seats from trace and other service usage. Phoenix’s free self-hosted software shifts the bill toward infrastructure and any external models used for evaluation. Those numbers are not interchangeable: a conversation can contain multiple recorded operations, and an evaluation can generate model charges outside the observability subscription.

Which LLM observability tool should you choose?

Choose Langfuse for a connected tracing and prompt workflow, LangSmith for a managed platform alongside agent development, or Phoenix for a self-hostable investigation environment. Each supports meaningful iteration; the best fit depends on the team’s stack and preferred operating model.

Sources & methodology

This article uses vendor documentation and our editorial analysis. We have not performed a controlled hands-on test. Product facts were checked on Sep 19, 2026; each profile records its own verification date.

How we evaluate tools · Suggest a correction

Explore the tools in this guide

Langfuse

An AI engineering platform connecting application traces, cost and latency metrics, prompt versions, datasets, and evaluation results.

Free optionCore: $29/month plus excess usage

LangSmith

An AI application platform for tracing agents, inspecting production behavior, gathering feedback, and turning recorded runs into evaluation datasets.

Free optionPlus: $39/seat/month plus usage

Arize Phoenix

A self-hostable AI observability workspace for inspecting traces, replaying model calls, managing prompts, and comparing experiments on shared datasets.

Free optionFree to self-host; infrastructure extra

Keep exploring

All articles
0 tools selected for comparison
LLM observability tools: Langfuse vs LangSmith vs Arize Phoenix | AIPicksy