Case studies and results from real engagements.

On October 6, 2026, New Relic announced AI Evaluation, a new capability inside its AI Observability product. The company says it scores the quality and safety of AI responses in the same distributed traces teams already use to debug slow or failing requests. It is due in public preview in November, so nobody can run it yet. Source
That timing is useful. The product does not exist for you to buy today, but the work it depends on does exist, and you can start it this week. If your agent traces do not record which model, prompt version and tool calls produced an answer, no evaluation product can fix that later. This article covers what New Relic claims, what the OpenTelemetry GenAI conventions already define, and a practical way to get ready. Everything here is based on the announcement and the published conventions; we have not tested the product.
New Relic describes AI Evaluation as a framework that looks at the whole application transaction rather than isolated single-model calls. It uses an asynchronous "LLM-as-a-judge" service to scan live telemetry and score problems such as hallucinations, prompt injections and data leaks. Those scores are attached as attributes to the distributed trace, so an engineer can see the prompt, the judge's reasoning and the surrounding application behavior on one screen. Source
The release lists several parts. Configurable guardrails check sampled inputs and outputs for prompt injection, jailbreak attempts, personal data leaks, toxicity and bias. For retrieval-augmented generation, it mentions faithfulness and answer relevancy, which helps separate a model's reasoning from a retrieval failure. It also says teams can compare quality against token cost to see whether an expensive model earns its price. Pre-built evaluators are included. Source
Before production, the release describes a prompt playground, versioned "golden datasets" built from real traces or synthetic data, and controlled experiments across prompts, models and configurations. Source
One framing from the release is worth keeping. New Relic says teams need to see which model responded, which prompt was used, what tools an agent called, how much the interaction cost and whether the result was useful. That list is a good checklist for your own telemetry, whichever tool you pick. Source
The announcement is a vendor press release, and it has gaps. It gives November as the public preview window but no price, no supported framework list, and no accuracy data for the judge. It does not say how often the judge is wrong, how sampling is chosen, or what an evaluation adds to your telemetry bill. Treat the claims as intent until the preview shows them working on real traffic. Source
An LLM judge is also a model. It can be wrong, and it can be biased toward long or confident answers. A score attached to a trace is a signal for a human to look at, not a verdict. Our recommendation is to sample judge results against human review for any workflow where a mistake costs money or trust.
Evaluation is only as good as the traces it sits on. The OpenTelemetry project maintains semantic conventions for generative AI, and their agent section defines span types you can use today. Operation names include invoke_agent, execute_tool, plan, create_agent and invoke_workflow. Source
That matters because an agent run is not one model call. It is a plan, several model calls and tool executions, sometimes inside a workflow. If each step is its own span under one trace, a quality score on the run can point to the step that went wrong. Without that structure, a bad answer is just a bad answer.
The conventions also define attributes for the facts New Relic says teams need. Examples include gen_ai.request.model, gen_ai.agent.name, gen_ai.agent.version, token usage such as gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, and error.type. There are also gen_ai.prompt.name and gen_ai.prompt.version for named prompt templates. Source
One caveat is important. The agent spans document is marked with a Development status, and many of these attributes are listed as Development too. Names can change. Wrap your instrumentation in one small module so a rename is a one-file change, and pin the version of any instrumentation library you adopt. Source
The same conventions list prompt and response content, such as input messages, output messages, system instructions and tool definitions, as Opt-In attributes. Opt-In means off by default, and for good reason. Prompts and answers often contain customer names, account numbers and internal documents. Source
An evaluation judge needs to read content to grade it, so there is a real tension. Turning content capture on everywhere gives the judge material, and it also copies sensitive text into your observability store. Decide this per workflow. For a low-risk internal summarizer, capture may be fine. For a support agent handling billing data, redact first, restrict who can read the traces, and set a short retention period. This is a design choice for your own security and privacy review, not something the announcement settles.
Here is a plan that works whether you end up using New Relic, another vendor, or your own scripts. It is our recommendation, not a requirement from any source above.
Day 1: pick one workflow. Choose an agent task where a wrong answer has a visible cost, such as a refund reply or a ticket triage decision. Write down in one sentence what "good" means for it.
Day 2: record the basics on every run. Model name, prompt name and version, agent version, tool names called, token counts and the final status. If you already emit OpenTelemetry traces, add these as span attributes using the GenAI names where they exist.
Day 3: shape the trace. Make sure one user request produces one trace, with the agent run as a parent span and each model call and tool call as children. Check that a failed tool call shows an error on its span.
Day 4: build a small golden set. Pull about fifty real requests with known-good outcomes from tickets or logs. Remove personal data first. The announcement describes versioned golden datasets, but the file itself is yours and will move between tools.
Day 5: score by hand. Grade twenty runs yourself on correctness, safety and whether the right action was taken. This gives you a baseline to compare any automated judge against later.
Day 6: add cost to the picture. Join token counts to your price list so each run has a cost. A model change is a price decision as much as a quality one, which is the point the release makes about quality against token cost. Source
Day 7: decide the guardrail rules. List the events that should page a human, such as a suspected prompt injection or a personal data leak in an answer. Decide who gets the alert before any tool is switched on.
If your team wants help with this kind of instrumentation, our AgentOps services cover monitoring and operating agents in production.
Evaluation scores answer a question about quality. They do not replace the other controls. Approval steps for consequential actions, least-privilege tool access, rate limits and rollback plans still matter, because a judge that runs after the fact cannot undo a payment or a deleted record.
A sensible order is to prevent what you can, observe everything else, and evaluate a sample. Traces tell you what happened. Evaluation tells you how well it went. Human review decides what to change. Each layer is weak without the others.
Our view is that the most valuable idea in the announcement is the simplest one: attach quality to the trace instead of keeping it in a separate tool. Whether New Relic's version delivers on that is a question for the November preview. The groundwork of clean traces, versioned prompts and a small golden set pays off with any option.
No. New Relic says it will be available in public preview in November. Until then, you can prepare traces, prompts and test data. New Relic
The press release does not say. It names no supported frameworks or standards, so ask New Relic before you plan around it. New Relic
Not yet. The agent spans document carries a Development status, so names and fields may change. Isolate your instrumentation in one module. OpenTelemetry GenAI
We would not rely on it alone. A judge is a model and can be wrong. Compare its scores with human grades on a sample, and keep human review for high-cost decisions.
Only after a privacy review. The OpenTelemetry conventions treat message content as Opt-In because it can contain sensitive data. Redact, limit access and set retention first. OpenTelemetry GenAI
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.