CIO Insider

CIOInsider India Magazine

Separator

New Relic Introduces AI Evaluation Across the Developer Lifecycle

CIO Insider Team | Friday, 9 October, 2026
Separator

New Relic, the intelligent observability platform, today announced a framework for advanced AI evaluation. New Relic AI Evaluation is a new capability within New Relic AI Observability.

Unlike point solutions that evaluate isolated single-LLM calls, it delivers transaction-level business impact insights across the full developer-to-production lifecycle, tracking response quality and behavior while providing automated, real-time insights into AI guardrail performance.

As generative AI becomes deeply embedded across enterprise workflows, software is shifting from deterministic code to probabilistic systems that fail in various ways, from subtle hallucinations to unexpected performance shifts when frontier vendors update LLMs.

This creates a new observability requirement as traditional application signals no longer tell the whole story. Teams now need to see which model responded, which prompt was used, what tools an agent called, how much the interaction cost, and whether the result was useful.

Observing AI in a silo or through best-of-breed point solutions creates dangerous blind spots around upstream and downstream operational impacts while widening the visibility gap between AI developers and production engineering.

A single, unified intelligent observability platform built for the AI era solves this by uniting full-stack operational telemetry with real-time AI evaluations—delivering the essential trust layer and end-to-end visibility leaders need to control costs, mitigate security risks, and safely scale enterprise AI.

Built natively into the unified New Relic platform, AI Evaluation goes beyond isolated prompt checks to analyze AI performance across the entire application lifecycle—down to the underlying transaction. By uniting full context rather than focusing on a single metric, it embeds quality and security guardrails directly into existing application performance monitoring workflows.

Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime

Replacing unscalable manual reviews with an asynchronous "LLM-as-a-judge" service, AI Evaluation automatically scans live telemetry to score vulnerabilities like hallucinations, prompt injections, and data leaks. It attaches these probabilistic quality scores as attributes directly to deterministic distributed traces, allowing teams to isolate the exact root cause of a failure, whether in the prompt, vector database, or backend infrastructure, within a single view. By linking qualitative response scores directly to underlying compute consumption, teams can easily determine if expensive models deliver enough semantic value over faster, lower-cost alternatives.

Also Read: AI Autonomy Race: How Advanced are Top Countries' AI Strategies?

New Relic Chief Product Officer Brian Emerson says, "Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime. Unlike point solutions that evaluate isolated LLM calls, we look holistically across the entire application transaction to show exactly what happens whenever AI is involved, giving teams full visibility into technical health, business impact, and performance—all within the platform tools that SREs, platform engineers, and developers already use.”



Current Issue
The Flight Of Aspiration: Uttarakhand Boy Ravi Tamta Builds Electric Flying Car



🍪 Do you like Cookies?

We use cookies to ensure you get the best experience on our website. Read more...