The verdict
Arize Phoenix appears in 15 AI-ranked categories — best position #2 for self-hosted llm observability tool.
Excellent OpenTelemetry/OpenInference foundation, broad Python/TypeScript/Java instrumentation, strong trace debugging, RAG analysis, datasets, experiments, and evaluations; near-tied with OpenLIT, ranking higher for its polished investigation and evaluation workflow.
Claude OTel-native tracing built on the OpenInference conventions, runs as a single container for near-zero-friction on-prem starts, and has the strongest open-source evaluation and dataset/experiment workflow of the group — near-tie with Langfuse for eval-centric teams, ranked second on weaker multi-tenant/production hardening
Grok Strong open-source/self-hostable (ELv2) observability & evaluation tool focused on LLM traces, RAG/agent debugging, evals (LLM-as-judge), and OpenInference/OTel standards; simple Docker/Python deployment keeps data fully on-prem; excels for experimentation-to-prod workflows in privacy-sensitive environments. Near-tie with Langfuse on core merits but edges on lighter ML-native eval depth for some teams.
Gemini OpenTelemetry-native (using OpenInference standards) and highly optimized for deep RAG analysis, model evaluations, and vector visualization, with a fantastic local notebook execution model.
Where Arize Phoenix falls short, per the models
- GPT Less complete as a turnkey multi-team operations platform for alerting, governance, and long-term fleet management.
- Claude Auth, RBAC, and multi-team production operation are much thinner than Langfuse; it shines for engineer-driven debugging and evals, not as a hardened shared platform service
- Gemini Lacks built-in prompt management, API gateways, and multi-tenant access controls in the open-source version, making it harder to use as a shared central enterprise portal.
- Grok Phoenix OSS is more dev/experimentation-oriented than full enterprise production platform (lacks some scaling/polish of Langfuse's ClickHouse backend).
Poll history — #2 in all 2 polls since Jul 13
#2 → #2
Top alternatives per the models: Langfuse · Helicone · OpenLLMetry · Opik
Open-source, OpenInference/OTel instrumentation with a strong, well-regarded eval library and excellent developer-to-production workflow for tracing, retrieval debugging, and drift; the paid Arize AX path extends it to enterprise-scale production monitoring. Near-tie with Langfuse.
Grok Single-process free self-host (pip/Docker), fully OpenTelemetry/OpenInference native for portable spans, excellent built-in evals/experiments/RAG introspection, no event caps, and notebook-to-production path; strong ML heritage makes it the cleanest standards-based choice when data control and zero platform cost matter most. Near-tie with Langfuse for pure tracing/eval depth.
GPT Strong open-source, OpenTelemetry-native tracing with broad OpenInference instrumentation, excellent RAG and agent evaluation, datasets, experiments, prompt tooling, and unrestricted air-gapped self-hosting
Gemini Open-source, vendor-agnostic, and strictly OpenTelemetry-native with standout capabilities in RAG retrieval inspection, embedding drift visualization, and notebook-to-production evaluation pipelines.
Where Arize Phoenix falls short, per the models
- GPT Teams needing mature managed alerting and large-scale production analytics may need Arize AX or additional infrastructure
- Claude The Phoenix-vs-Arize-AX split (free dev tool vs paid platform) is confusing, and Phoenix alone is lighter on production-scale governance, alerting, and prompt management.
- Gemini Prompt lifecycle management and production team-collaboration features are sparse compared to Langfuse/LangSmith; requires more manual infrastructure setup for production scale.
- Grok Elastic License (source-available, not full OSI) and lighter product-facing prompt/session/ops UI than Langfuse; high-scale production features push toward paid Arize AX.
Poll history — On this board 10 of 10 polls since Jun 29 · now #2
#3 → #4 → #3 → #4 → #3 → #4 → #3 → #3 → #3 → #2
What changed in the models’ minds
ClaudeJul 16 → Aug 15 poll
- Newretrieval debugging
- NewNear-tie with Langfuse“Near-tie with Langfuse.”
- Newproduction-scale governance and alerting“production-scale governance, alerting”
- Droppedembeddings“embeddings/drift analysis”
+2 more changes
GeminiJul 16 → Aug 15 poll
- Newvendor-agnostic
- NewPrompt lifecycle and team-collaboration features are sparse“Prompt lifecycle management and production team-collaboration features are sparse compared to Langfuse/LangSmith”
- Newmanual infrastructure setup for production scale“requires more manual infrastructure setup for production scale.”
- DroppedJupyter Notebooks and data science analysis“The UI and workflow are heavily designed for Jupyter Notebooks and data science analysis rather than production application developer tracing.”
GrokJul 16 → Aug 15 poll
- Newsource-available not full OSI“Elastic License (source-available, not full OSI)”
- Newlighter product-facing prompt/session/ops UI“lighter product-facing prompt/session/ops UI than Langfuse”
- Newhigh-scale production features push toward paid Arize AX
- Droppedopen-source“Strong open-source (ELv2)”
+2 more changes
Top alternatives per the models: Langfuse · LangSmith · Braintrust · Helicone
Strongest open-source specialist for tool calling, with separate Tool Selection and Tool Invocation evaluators covering wrong-tool, wrong-argument, parallel-call, and no-call cases, plus OpenTelemetry-native traces and experiments. Near-tied with LangSmith; it ranks higher for accessibility and purpose-built metrics.
Grok Source-available with explicit ToolSelectionEvaluator, ToolInvocationEvaluator and ToolResponseHandlingEvaluator plus ordered-trajectory and path-convergence scoring; OTel-native self-hosting with no event caps delivers concrete tool-call reliability measurement tied to full traces at low operational cost
Gemini Open-source and OpenTelemetry-native architecture offering zero vendor lock-in, deep span-level inspection of tool execution paths, and self-hosted privacy for enterprise agent evaluations.
Claude Strongest open-source option — OTel-based span tracing surfaces tool-call spans cleanly, with prebuilt tool-calling eval templates and the ability to self-host, making it the best free/private choice for inspecting and grading tool trajectories.
Where Arize Phoenix falls short, per the models
- GPT Its semantic tool evaluators rely heavily on LLM judges and self-hosting requires operational work; it is not a complete deterministic agent-simulation harness.
- Claude Its evaluators lean on LLM-as-judge templates you must tune, and serious scale/collaboration features push you toward the commercial Arize AX tier.
- Gemini Requires self-hosting setup and ongoing infrastructure management for long-term trace storage and custom metric dashboards.
- Grok Not for pure no-code teams or those unwilling to operate a self-hosted instance
Poll history — On this board 2 of 2 polls since Aug 3 · now #2
#3 → #2
Top alternatives per the models: Braintrust · LangSmith · DeepEval · Galileo
The strongest open-source choice for LLM/GenAI observability — OpenTelemetry-native tracing, span-level inspection, built-in and LLM-as-judge evals, and embedding/retrieval analysis, all fully self-hostable with no vendor lock-in. Best pick for teams whose "models" are LLM apps, RAG, or agents.
Gemini OpenTelemetry-native open-source framework delivering deep tracing, embedding visualization, and evaluation metrics for both LLM apps and traditional ML models, eliminating vendor lock-in; near-tie with Evidently AI assuming microservice architectures.
Grok Strongest self-hostable option for embedding-heavy, multimodal, and LLM/agent workloads with OpenTelemetry/OpenInference-native tracing, trajectory evals, and drift visualization; runs locally or in Docker with low friction; shares evaluation logic with production-grade Arize AX while staying free for internal use
Where Arize Phoenix falls short, per the models
- Claude Purpose-built for LLM/embedding observability; it is not the tool for monitoring classical tabular model accuracy, feature drift, or data-quality at scale.
- Gemini Enterprise capabilities like automated anomaly alerting, role-based access controls, and long-term historical root-cause analysis require subscribing to the proprietary Arize SaaS platform.
- Grok Elastic License 2.0 (source-available, not pure OSI open-source) restricts offering it as a service; live production alerting and advanced fleet features sit in the paid AX tier
Poll history — On this board 2 of 2 polls since Aug 3 · now #3
#2 → #3
Top alternatives per the models: Evidently · NannyML · whylogs · Alibi Detect
Best when the real problem is diagnosing why retrieval fails — OpenTelemetry-based tracing ties each generation back to retrieved chunks, with span-level LLM-as-judge evals (hallucination, relevance, QA correctness) and embedding/retrieval visualizations; open-source and self-hostable with a clean local dev loop.
GPT Best free self-hosted all-in-one workflow: OpenTelemetry/OpenInference tracing exposes retrieval and generation steps, while RAG evaluators, datasets, experiments, human labels, evaluator traces, and prompt replay make failures actionable.
Gemini OpenTelemetry-native platform that seamlessly links span-level pipeline tracing with automated RAG evaluations (retrieval relevance, groundedness, hallucination detection), excelling at root-cause debugging of chunking and retrieval failures.
Grok OpenTelemetry-native tracing captures every retrieval span and chunk so you can visually inspect exactly why a retrieved context failed and why the answer went wrong, paired with solid built-in RAG relevance/faithfulness evaluators and versioned datasets/experiments; the strongest debugging and iteration loop for practitioners who have already instrumented their pipeline.
Where Arize Phoenix falls short, per the models
- GPT The full observability platform is overkill for teams needing only lightweight local regression tests.
- Claude The tracing/observability surface is heavier than teams who just want a score want to adopt, and its judge templates are less RAG-metric-opinionated than Ragas, so you assemble the metric suite yourself.
- Gemini Requires deeper instrumentation and deployment overhead than pure Python scoring libraries, which is excessive for teams only needing quick offline dataset evaluations.
- Grok It is an observability platform first, so pure offline metric computation without tracing overhead is more friction than a library.
Poll history — #3 in all 6 polls since Jul 11
#3 → #3 → #3 → #3 → #3 → #3
Top alternatives per the models: Ragas · DeepEval · TruLens · Opik
OpenTelemetry- and OpenInference-native, making it the most vendor-neutral instrumentation path; excellent evals (including hallucination/retrieval/agent-specific templates), embedding/drift analysis, and a free OSS Phoenix tier that scales up to the enterprise AX platform.
Gemini Open-source, OpenInference/OTel standard-native agent tracing that uniquely bridges step-level tool execution debugging with embedding analysis, drift detection, and automated eval benchmarks.
Grok Strongest pure OSS diagnostic tool for agent trajectories via OpenInference + OTel; excellent embedding/drift analysis and eval library that surfaces silent degradation; single-container self-host makes it the lightest serious option for local-to-prod debugging of multi-agent flows
GPT Strong open-source choice for privacy and standards-first teams: OpenTelemetry/OpenInference instrumentation, broad agent-framework coverage, trace and session inspection, evaluations, datasets, experiments, and unrestricted self-hosting.
Where Arize Phoenix falls short, per the models
- GPT Phoenix itself lacks the richer continuous online evaluation, threshold alerting, and operational monitoring of Arize AX, so it is weaker as a complete production control room.
- Claude Phoenix self-hosted carries an ops burden and its UI is more analysis- than debugging-oriented; the polished, scalable experience lives in paid AX, which is enterprise-priced.
- Gemini Lacks native prompt versioning and developer playground features, making it better for post-hoc telemetry/eval analysis than day-to-day prompt/agent iteration.
- Grok Source-available (ELv2) rather than fully permissive OSS; oriented toward ML/data teams so the full enterprise AX path adds cost and complexity for pure
Poll history — On this board 5 of 5 polls since Jul 12 · #3 the last 4
#4 → #3 → #3 → #3 → #3
What changed in the models’ minds
GrokJul 12 → Aug 14 poll
- Newsingle-container self-host“single-container self-host makes it the lightest serious option for local-to-prod debugging of multi-agent flows”
- NewSource-available rather than fully permissive OSS“Source-available (ELv2) rather than fully permissive OSS”
- Newfull enterprise AX path adds cost and complexity“the full enterprise AX path adds cost and complexity”
- Droppedunified ML+LLM monitoring
+1 more change
GPTJul 15 → Aug 14 poll
- Newdatasets
- Newcontinuous evaluation, alerting, and monitoring gaps“Phoenix itself lacks the richer continuous online evaluation, threshold alerting, and operational monitoring of Arize AX”
- Droppedlanguage and provider integrations“broad Python, TypeScript, Java, provider, and agent-framework integrations”
- Droppedsetup and observability expertise“Phoenix requires more setup and observability expertise than polished managed products, particularly for large production deployments.”
ClaudeJul 14 → Jul 15 poll
- Newagent trace visualization“strong agent trace visualization”
- NewLangfuse near-tie“near-tie with Langfuse”
- Newresearcher-oriented UX“its UX is more researcher-oriented than ops-oriented”
- Droppedvendor-neutral instrumentation“vendor-neutral instrumentation means no lock-in”
+2 more changes
Top alternatives per the models: Langfuse · LangSmith · Braintrust · Datadog LLM Observability
Near-tie with Langfuse for teams prioritizing open standards and evaluation depth; excellent OpenTelemetry-native tracing, agent graphs, experiments, prompt iteration, and evaluators, with a managed free tier covering 25k spans monthly and an open-source local option.
Grok Completely free self-host (local/Docker, under 30 min, no limits), open-source core, excellent for RAG/debugging with low overhead once running, OTel-native for minimal custom code in compatible setups.
Claude Fully open-source and free with no gating — pip install and it runs locally or in a notebook, built natively on OpenTelemetry/OpenInference so instrumentation is standard and portable; best-in-class trace visualization and eval tooling for the price of zero.
Gemini The gold standard for zero-ops local debugging and evaluations, running via a simple terminal command (pip install) inside Jupyter notebooks or local servers. Natively built on OpenTelemetry standards, allowing easy telemetry migrations without vendor lock-in.
Where Arize Phoenix falls short, per the models
- GPT The managed free tier is single-developer, limited to 1 GB and 15-day retention, while self-hosting sacrifices the minimal-ops advantage.
- Claude Minimal-ops it is not once you leave the laptop — persistent team deployments mean self-hosting it yourself (or moving to paid Arize AX), so it suits teams comfortable running a container over those wanting a managed dashboard.
- Gemini Lacks a managed, startup-friendly cloud free tier for production logging, forcing teams to choose between the operational overhead of self-hosting or upgrading to the expensive enterprise-oriented Arize Cloud.
Poll history — #3 in all 2 polls since Jul 13
#3 → #3
Top alternatives per the models: Langfuse · Helicone · LangSmith · Braintrust
Best open-source-first choice for tracing and diagnosing heterogeneous agents, with OpenTelemetry/OpenInference interoperability, datasets, experiments, span and trace evaluation, and explicit trajectory evaluation over ordered tool calls. Particularly strong when observability and root-cause analysis matter as much as pass/fail scores.
Grok Robust open-source/self-hosted OTel-native tracing with solid trajectory and path-convergence evals, agent function-calling metrics, and ML heritage for production observability; strong for teams prioritizing data control, custom setups, and extending existing monitoring without lock-in. FIX: Narrower agent-specific metric depth and multi-turn simulation compared to top options; requires more setup for full custom evaluator pipelines.
Claude Excellent OSS tracing-plus-evals combo built natively on OpenTelemetry/OpenInference; strong prebuilt eval templates (hallucination, tool-call correctness, agent trajectory) that run locally at scale, and a clean path from notebook experimentation to the commercial Arize AX platform for production; near-tie with Langfuse — Phoenix has better eval libraries, Langfuse better product polish and prompt management.
Gemini Strongest OpenTelemetry-native, open-source tracing tool for session-level debugging. It excels at local execution, trajectory tracking, and native integrations with LlamaIndex and custom python agent frameworks.
Where Arize Phoenix falls short, per the models
- GPT Building sophisticated agent-specific test harnesses and trajectory rubrics still requires more hands-on assembly than the leading commercial platforms.
- Claude The OSS/commercial split is confusing and the self-hosted product is less turnkey for team collaboration (auth, multi-project) than Langfuse — expect to graduate to paid Arize for serious production use.
- Gemini Built primarily as an observability tool, making it less optimized for developer-focused CI/CD regression testing and metric playgrounds.
Top alternatives per the models: LangSmith · Braintrust · DeepEval · Langfuse
The strongest open-source option for teams standardizing on OpenTelemetry — OpenInference semantic conventions, excellent embedding/drift visualizations inherited from Arize's ML-monitoring DNA, strong retrieval/RAG evaluation tooling, and a clean upgrade path to Arize AX for enterprise scale; runs locally in a notebook in one line, which no competitor matches for experimentation-to-production continuity. Near-tie with Langfuse — Phoenix wins on evals and OTel purity, Langfuse on prompt management and all-in-one polish.
Grok Strong open-source (ELv2) observability with notebook-first workflows, robust RAG/agent evaluation templates (hallucination, relevance, toxicity), OpenTelemetry integration, backed by Arize's ML monitoring expertise for production-scale drift and quality monitoring in high-volume environments.
Gemini Built from the ground up on OpenTelemetry and the OpenInference standard, ensuring complete data portability and zero vendor lock-in. It allows enterprise machine learning teams to seamlessly unify LLM observability with existing data platforms, and its open-source library provides top-tier local troubleshooting and embedding visualization.
Where Arize Phoenix falls short, per the models
- Claude The OSS Phoenix product and the commercial Arize AX platform are distinct enough that teams outgrowing Phoenix face a real migration, and the UI is less refined for non-technical stakeholders reviewing traces.
- Gemini Setting up and hosting Phoenix at scale requires significant operational overhead, and its user interface is designed for data scientists and ML engineers, which can feel overly complex for application developers.
- Grok Phoenix open-source is more dev/notebook-oriented; full enterprise features (deeper production scaling, compliance) push toward paid Arize AX (not the simplest for lightweight proxy-style monitoring).
Poll history — On this board 2 of 2 polls since Jul 18 · now #2
#4 → #2
Top alternatives per the models: Langfuse · LangSmith · Braintrust · Datadog LLM Observability
Strong OTel-native observability and tracing for production debugging/RAG/agent monitoring, open-source core with good LLM-as-judge and drift detection; self-hostable option appeals to teams prioritizing control and standards compliance.
Claude Best OpenTelemetry-native open-source option — OpenInference tracing, a solid LLM-as-judge eval library, and dataset/experiment tracking that runs locally or in a notebook for free, with a clean upgrade path to Arize AX for enterprise scale.
Gemini The strongest open-source, OpenTelemetry-native framework for deep span-level tracing and embedding analysis. It is a near-tie with Langfuse but ranks slightly lower because it targets data scientists over typical developers.
Where Arize Phoenix falls short, per the models
- Claude The OSS product's collaboration, alerting, and hosted-team features are thin — multi-user production monitoring effectively pushes you into the paid Arize platform.
- Gemini Its engineering-heavy interface lacks collaborative workflows for non-technical stakeholders like product managers.
- Grok Less depth in full CI/CD eval gating or broad non-RAG agent simulation compared to top integrated platforms.
Poll history — #4 in all 3 polls since Jul 11
#4 → #4 → #4
What changed in the models’ minds
ClaudeJul 12 → Jul 13 poll
- NewLLM-as-judge eval library“a solid LLM-as-judge eval library”
- Newdataset/experiment tracking“dataset/experiment tracking that runs locally or in a notebook for free”
- Newalerting features are thin“alerting, and hosted-team features are thin”
- DroppedML-observability depth“Arize's ML-observability depth (drift, embeddings analysis)”
+2 more changes
GeminiJul 12 → Jul 13 poll
- Newstrongest open-source framework“The strongest open-source, OpenTelemetry-native framework”
- Newnear-tie with Langfuse“It is a near-tie with Langfuse”
- Newtargets data scientists“it targets data scientists over typical developers”
- Droppedstatistical ML monitoring capabilities“advanced statistical ML monitoring capabilities”
+1 more change
Top alternatives per the models: Braintrust · LangSmith · Langfuse · MLflow
Strongest open-source all-in-one option, combining OpenTelemetry-based tracing, datasets, experiments, annotations, prompt iteration, and pluggable evaluators while remaining framework- and model-neutral
Claude Strongest open-source, self-hostable option — OpenTelemetry-native tracing plus a good library of pre-built evaluators (hallucination, relevance, RAG); runs locally in a notebook or as a service with no vendor lock-in, and pairs with Arize AX if you later need production monitoring.
Where Arize Phoenix falls short, per the models
- GPT Self-hosting and operating it requires more infrastructure effort than using a polished managed service
- Claude Eval UX and managed-workflow polish trail the commercial leaders; you'll do more assembly, and heavy production observability pushes you toward the paid Arize tier.
Poll history — On this board 9 of 10 polls since Jun 29 · now #4
#3 → #3 → #5 → #3 → #3 → #4 → – → #7 → #6 → #4
What changed in the models’ minds
GPTJul 14 → Jul 15 poll
- Newprompt iteration
- Newframework- and model-neutral
- Newmanaged service comparison“than using a polished managed service”
- Droppedhuman labels
+2 more changes
Top alternatives per the models: Braintrust · DeepEval · LangSmith · Promptfoo
Strong open-source, OpenTelemetry/OpenInference-native choice for inspecting complex agent traces and evaluating tool selection, parameters, planning, path convergence, and whole trajectories; broad framework interoperability materially improves its value.
Claude Open-source tracing and evals built on OpenInference/OTel standards, solid agent-specific evals (tool-choice, path convergence), notebook-friendly for experimentation, with a credible enterprise upgrade path via Arize AX.
Grok Open-source observability with strong OTel support, agent evaluators, production monitoring, and self-hosting option; practical for tracing tool use and trajectories in diverse/multi-framework environments at lower cost
Where Arize Phoenix falls short, per the models
- GPT It remains more observability-and-analysis-centric than a turnkey regression-testing system, so dataset operations and CI workflows can require extra assembly.
- Claude The OSS product is more an observability-plus-evals library than a full managed platform — teams wanting hosted collaboration, RBAC, and dataset workflows out of the box must step up to paid Arize.
- Grok Less specialized depth in agent-specific step-level scoring or CI/CD eval automation compared to leaders; more general ML-focused
Poll history — On this board 2 of 2 polls since Jul 13 · now #4
#5 → #4
Top alternatives per the models: Braintrust · LangSmith · DeepEval · Langfuse
Open-source OTel-native observability and evaluation with strong drift detection, tracing, and troubleshooting suited for production ML/agent monitoring; self-hostable with rigorous evaluation harness.
GPT Best open-source debugging-focused alternative: OpenTelemetry-native traces, broad framework support, experiments, dataset evaluation, LLM judges, and excellent visibility into retrieval, tool calls, latency, and agent paths.
Where Arize Phoenix falls short, per the models
- GPT It is stronger at observability and diagnosis than at rich behavioral simulation or release-test orchestration.
- Grok Heavier on traditional ML observability than lightweight agent-specific tracing or rapid experiment workflows for some dev teams.
Poll history — On this board 2 of 2 polls since Jul 14 · now #4
#7 → #4
Top alternatives per the models: LangSmith · Braintrust · Langfuse · Maxim AI
Best near-tie for open-source, self-hosted evaluation workflows: it combines versioned and tagged prompts with a capable playground, datasets, experiments, tracing, and Python and TypeScript clients.
Where Arize Phoenix falls short, per the models
- GPT Its prompt-management safeguards and production retrieval workflow are less mature, requiring careful caching and fallback design.
Poll history — On this board 2 of 10 polls since Jul 14 — off it in the latest
– → – → – → – → – → – → – → #7 → #7 → –
What changed in the models’ minds
GPTJul 14 → Jul 15 poll
- Newtagged prompts
- NewPython and TypeScript clients
- Newcaching and fallback design“requiring careful caching and fallback design”
- Droppedmulti-provider playground
+2 more changes
Top alternatives per the models: Langfuse · LangSmith · PromptLayer · Braintrust
OpenTelemetry-native observability and evaluation tool that fits seamlessly into modern enterprise infrastructure without vendor lock-in.
Grok Robust OTel-native tracing and visualization for embeddings/hallucinations with production monitoring suitable for complex apps
Where Arize Phoenix falls short, per the models
- Gemini Streamline the developer experience for quick local setups and basic unit-testing without requiring full telemetry pipeline configuration.
- Grok Enhance developer-centric testing framework and pytest-style integration for faster iteration in code-first workflows
Poll history — On this board 1 of 2 polls since Jul 12 — off it in the latest
#6 → –
Top alternatives per the models: DeepEval · Promptfoo · Ragas · Inspect AI
Head-to-head — how the models call it
Watch Arize Phoenix
Boards re-poll weekly and the models change their minds. One short email only when Arize Phoenix's standing moves — a rank change, a rival overtaking, or new reasoning from the models. Nothing otherwise.
Embed your ranking badge
Arize Phoenix ranks #2 for best self-hosted llm observability tool by AI-model consensus. Put the badge in your README, docs or site — it updates automatically as the models re-rank.
[](https://modelsagree.com/best/best-self-hosted-llm-observability?utm_source=badge&utm_medium=embed&utm_campaign=badge-arize-phoenix)<a href="https://modelsagree.com/best/best-self-hosted-llm-observability?utm_source=badge&utm_medium=embed&utm_campaign=badge-arize-phoenix"><img src="https://modelsagree.com/badge/arize-phoenix.svg" alt="Arize Phoenix — ranked #2 for Best self-hosted LLM observability tool by AI models on ModelsAgree" height="28"></a>Rankings are computed from what the models answer, re-polled on demand · raw reasoning shown verbatim · methodology