Skip to main content
Turn production failures into evaluations, tested prompt improvements, and custom views built for your team. Alyx works with your traces, annotations, datasets, and experiments in Arize AX to plan and carry out multi-step tasks. Alyx uses your current context, asks for missing information, and presents proposed changes for review. Your conversation follows you across the platform, so you can investigate a failure and test an improvement without rebuilding context at each step.

What can Alyx do?

Turn failures into evaluation coverage

“Investigate this project’s recurring errors and build an evaluation to catch one of the failure patterns.”
Start with a project that has traces. Alyx finds relevant examples, inspects inputs, outputs, and tool calls, and groups related failures. It can then propose an evaluator and configure its evaluation task, including data sources and variable mappings. Result: a diagnosis backed by trace evidence and an evaluation you can reuse on incoming data. New to evaluations? Start with the AI agent evaluation handbook.

Alyx can create an eval for you directly from your traces so that you can start measuring what matters without any manual setup.

Compare prompts with evidence

“Using this dataset, create two prompt variants, attach a helpfulness evaluation, and compare their results.”
Start with a dataset and a configured model integration for the experiment. Alyx prepares the playground, creates variants, runs experiments, and analyzes results. You can also ask it to generate synthetic examples, refine the prompt, or save a version to Prompt Hub. Result: tested alternatives and results that help you choose your next change.
Comparing Prompts experiment view with Alyx sidebar suggesting prompt improvements to score better with experiments.

Alyx can help you iterate on improvements based on experiment results.

Build custom views for your application

“Create a trace view showing each tool call beside its response, latency, and evaluation scores.”
Open a trace, session, or annotation-queue record and describe the layout you need. Alyx generates a custom view against your data. Review it and request refinements before accepting it. Use the view controls to save and organize supported views for reuse. Result: a layout tailored to your debugging or review task, without exporting data or building a separate interface. Learn more about custom views.

Align an evaluator with human feedback

“Help align this evaluator with the human annotations in my dataset.”
Start with an evaluator and labeled examples. Alyx helps select the annotation column, configures an agreement check, runs experiments, and iterates on the evaluator template. If you need review data first, it can help create annotation configurations and labeling queues. Result: a revised evaluator and experiment results showing its agreement with your labels.

Explore all capabilities

Expand a category for supported actions and an example request. Some actions need additional context, a configured integration, or your acceptance before they run.
Generate and refine custom React views for traces and spans, sessions, or labeling-queue records. Inspect the underlying data, review the proposed layout, and use Customize Tabs to organize accepted views. Publishing a view to everyone requires being its creator and an org admin. Labeling-queue views can also be shown on all queues by their creator or an org admin. Queue views present record data; reviewers submit labels through the standard annotation panel.Try: “Build a session view showing the conversation and tool calls in order.”
Read issues recorded by Signal, inspect an individual issue, and follow its evidence into traces. An open Signal issue is attached as context. Signal performs the scheduled analysis; use Alyx to ask follow-up questions and investigate its findings.Try: “What issues has Signal found for this project? Help me investigate one.”
List traces; inspect span hierarchy, inputs, outputs, attributes, and errors; search by content or metadata; build semantic and multi-span filters; inspect the active filter and time range; categorize recurring failures or user requests. Open filtered trace links to inspect the evidence.Try: “Find traces where the search tool failed, then group the error messages.”
Compute counts, averages, sums, percentages, latency, cost, and token usage across spans or traces. Group by a dimension, sort results, and return top-N groups. List, create, and update custom metrics for reusable measurements.Try: “Show the five models with the highest total cost this week.”
List and inspect evaluators; suggest templates; create or update LLM-as-a-judge and code evaluators; configure classification choices and data sources; fix variable mappings; align evaluator templates with human annotations through experiments.Try: “Build a code evaluator that checks whether the response is valid JSON.”
Create and update project or dataset tasks; choose evaluators, filters, sampling, and run settings; configure trace/session multi-span queries and per-variable mappings. Run historical evaluations, set how many spans to evaluate and whether to skip previously evaluated spans, check status, and open task logs.Try: “Configure this session evaluation and preview the data its variables will receive.”
For Enterprise accounts with managed agents enabled, Alyx can propose agent-as-a-judge evaluators and project tasks. These require an Anthropic integration. They are separate from ordinary LLM and code evaluators, and task-editing options are more limited.Try: “Create an agent-as-a-judge evaluator for this project using my rubric.”
List and preview datasets, inspect columns and attached evaluations, filter dataset or experiment rows, create datasets from spans or synthetic examples, and append spans or generated examples to an existing dataset.Try: “Add these failure spans to my regression-test dataset.”
List, read, and load prompts; create or edit playground variants; optimize prompts against feedback; save prompts and new versions to Prompt Hub. Attach datasets, select columns, attach or detach evaluations, set run size, run experiments including background runs, check progress, and compare scores, outputs, and token usage.Try: “Run these variants on 50 rows and compare them with the previous experiment.”
Inspect annotation configurations; create labels and scores; annotate spans or dataset examples. List, create, and edit labeling queues. Review the proposed queue form and complete annotators and reviewer instructions before accepting.Try: “Create a labeling queue for reviewing failed responses from this project.”
Create and edit dashboards and their statistic, bar-chart, line-chart, text, and pivot-table widgets. Adjust metrics, filters, granularity, axis ranges, titles, and colors.Try: “Create a dashboard with p95 latency by model and total token usage.”
Find projects and saved assets, navigate to relevant AX pages with context, get API-key setup assistance, and search documentation. Get guidance on installing Arize coding-agent skills, instrumentation, and the AX CLI. Documentation assistance uses RunLLM with your consent.Try: “Help me set up Arize tracing with my coding agent.”

Where to find Alyx

Alyx home screen with greeting, chat input, model selector, and suggested actions including quickstarts for tracing, playground, and evaluators

Alyx home: ask a question, use @ for context, and try suggested quickstarts

Alyx is a built-in AI agent that works wherever you are in the platform. Your conversation carries over between pages and tasks, so you can go from a trace to the playground to an experiment without losing context. Alyx appears in two forms that share the same conversation:
  • Home view: the full-page experience on the Arize AX home page. Use it to start or resume a conversation, browse suggested workflows, and pick up a previous thread from your history. It is the place to begin broad or return to earlier work.
  • Side chat: a panel you open over any page with the keyboard shortcut. It inherits the context of whatever you are viewing, such as a project, trace, dataset, or experiment, so you can ask about the data in front of you and let Alyx act on it without leaving the page. Dock it to the right as a resizable sidebar or detach it as a floating window. When you open it from a trace, it scopes to that trace and its spans.
Each surface loads the skills relevant to the task at hand:

Open Alyx

Use the keyboard shortcut to open or close the side chat from anywhere in the app.
  • macOS: Cmd+L
  • Windows / Linux: Ctrl+L
You can customize this shortcut in Settings. Set any combination of a modifier key (Cmd, Ctrl, Alt, or Shift) and a second key, and Alyx saves it for future sessions in that browser.
Alyx Finding key takeaways from experiments on a dataset.

Alyx summarizing experiments.

More places to work with Alyx

  • Sessions: inspect interactions across traces and generate custom session views.
  • Labeling queues: propose queue configurations and generate custom views for reviewers.
  • Dashboards: create or edit widgets while reviewing the dashboard.
  • Signal: ask about a recorded issue with its context attached to your message.

Configure Alyx

Choose a model

Use the model selector in Alyx chat to choose an available Arize-managed model or a model from your own integration. Availability depends on your account and space configuration. Alyx filters for models with sufficient tool-use capability, so not every model configured for another AX feature appears in this selector.

Use your own integrations with Alyx

You can run Alyx on your own LLM providers. Configure them in Settings, then Account Settings, then Integrations, and select a configured model from the model selector in the Alyx chat. For setup details see AI provider integrations. Available integrations include the following providers. Model eligibility also applies; configuring a provider does not make every model available to Alyx:
  • OpenAI
  • Anthropic
  • Gemini
  • Azure OpenAI
  • Vertex AI
  • AWS Bedrock
  • NVIDIA NIM
  • Custom endpoints that expose an OpenAI-compatible API

Add and manage LLM integrations for Alyx in Settings

Add context to Alyx

Alyx works best when it knows exactly what you’re looking at. There are three ways to give it context:
  • Highlight: Select any text on the page, such as a span attribute, an error message, or a prompt snippet, then press Cmd+L on macOS or Ctrl+L on Windows and Linux. Alyx opens and adds the selected text to your message. If nothing is selected, the shortcut opens or closes Alyx.
  • Mention: Type @ in the Alyx input to open a menu. Mention a dataset, experiment, project, or span, and Alyx receives the IDs so it can scope the conversation to that data.
  • Type it: Include context directly in your message. Reference a trace, dataset, or experiment by ID, or start with “Additional context:” and add what Alyx needs to know.
Whatever you add is sent with your message as input context and displayed as pills in the Alyx input field. Alyx uses this context to scope its response to the data you care about, so you don’t have to restate the situation every time.

Settings

Open Settings from the menu in the top-right corner of the Alyx chat to personalize Alyx and configure approvals, your keyboard shortcut, documentation assistance, and tips. Keyboard shortcut: Change the shortcut used to open and close Alyx, as described in Open Alyx. Auto Accept: Alyx presents proposals for you to review before applying changes covered by the categories below. Turn on Auto Accept for a category to let Alyx apply those changes without a confirmation step. Each category is controlled independently:
  • Eval and task updates: Apply evaluator and task configuration changes automatically.
  • Annotations: Apply annotation configs and span annotations automatically.
  • Dataset creation and appends: Create datasets and append rows automatically.
  • Prompt changes: Apply prompt edits in the playground automatically.
  • Experiment runs: Start proposed experiment runs automatically.
Other actions, including labeling-queue and custom-view proposals, have their own acceptance steps. Auto Accept is not a blanket approval for every action. Tips: Show or hide product tips above the chat input while Alyx works. Documentation: Authorize the documentation support skill. When enabled, questions about Arize documentation are answered with help from RunLLM, as described in Third-party integrations. Personalize Alyx: Add personal context so Alyx can tailor its responses across chats. In Settings, you can configure:
  • Name: What Alyx should call you.
  • Role: What best describes your work.
  • Instructions: Free-form context Alyx should keep in mind, such as “I primarily code in Python.”
This context is per-user and is remembered across Alyx chats. Instructions are validated for safety before they are saved and support approximately 8,000 characters. If you have not configured personal context, Alyx may show a setup prompt in new chats.

Open Alyx Settings from the chat menu to configure your keyboard shortcut and Auto Accept settings

Chat experience

Edit and resend a message

Hover over a message you have sent and select Edit & resend, or click the message to edit it in place. Change the text and adjust the attached context, then press Enter to send or Escape to cancel. If that message previously produced results, such as a dataset, evaluation, or experiment, Alyx asks you to confirm before resending, because resending rewrites the conversation from that point on. Everything after the edited message is replaced, the turns above it stay in place, and the original turn is archived.

Queue a message

If Alyx is still responding, you can line up your next message instead of waiting. Type it and send, and it appears as Queued beneath the conversation. Alyx sends it automatically once the current response finishes. Send it right away with Send Now, or remove it from the queue.

Follow Alyx’s progress

For multi-step work, Alyx can display a plan and update it as it completes tasks. Supported Claude and GPT models also provide reasoning summaries in the thinking view.

Return to a conversation

Open your conversation history from Alyx home to resume earlier work. The home view and side chat share the same conversation as you navigate between pages.

Data privacy

Alyx sends requests to the model provider selected for your conversation. Arize-managed options include Azure OpenAI and Anthropic. When you use your own integration, your provider configuration and its data-handling terms apply. Data processing: Azure OpenAI and Anthropic act as data processors for prompts and outputs sent to and generated by Alyx, depending on which Arize-hosted model is used. Provider processing and retention depend on the service and applicable configuration; see the provider documentation linked below. Model training: Azure OpenAI and Anthropic describe their training and data-use commitments in the provider documentation linked below. Review the applicable terms for your selected integration. Microsoft and Azure OpenAI: The Azure OpenAI Service is fully controlled by Microsoft and hosted in Microsoft’s Azure environment. It does not interact with any other OpenAI-operated services such as ChatGPT or the OpenAI API. Anthropic and Claude: When Alyx uses Claude, Anthropic processes requests on Anthropic infrastructure under Anthropic’s security and data commitments for enterprise API use, separate from unrelated consumer products. Security and compliance: Azure OpenAI and Anthropic help Arize meet industry-standard security and compliance measures throughout the process. Diagram of Alyx data flow: customer data processed through Arize-hosted Azure OpenAI or Anthropic without exposure to unrelated third-party consumer services For more detail see Azure OpenAI Service data, privacy, and security in Microsoft’s documentation, Anthropic’s trust and privacy documentation, or contact support@arize.com.

Third-party integrations

Alyx includes a support skill that answers user questions. When you ask a support-related question, that question is sent to RunLLM for processing. Only the specific question you ask is shared with RunLLM. No additional model information or user data is included. You can manage authorization in Alyx Settings → Documentation. Contact support@arize.com if you need help. Before using the support skill for the first time you will be asked to acknowledge a one-time disclaimer outlining RunLLM’s involvement.