New Relic Introduces AI Evaluation Across the Developer Lifecycle
New Relic, the intelligent observability platform, today announced a framework for advanced AI evaluation. New Relic AI Evaluation is a new capability within New Relic AI Observability.
Unlike point solutions that evaluate isolated single-LLM calls, it delivers transaction-level business impact insights across the full developer-to-production lifecycle, tracking response quality and behavior while providing automated, real-time insights into AI guardrail performance.
As generative AI becomes deeply embedded across enterprise workflows, software is shifting from deterministic code to probabilistic systems that fail in various ways, from subtle hallucinations to unexpected performance shifts when frontier vendors update LLMs.
This creates a new observability requirement as traditional application signals no longer tell the whole story. Teams now need to see which model responded, which prompt was used, what tools an agent called, how much the interaction cost, and whether the result was useful.
Observing AI in a silo or through best-of-breed point solutions creates dangerous blind spots around upstream and downstream operational impacts while widening the visibility gap between AI developers and production engineering.
A single, unified intelligent observability platform built for the AI era solves this by uniting full-stack operational telemetry with real-time AI evaluations—delivering the essential trust layer and end-to-end visibility leaders need to control costs, mitigate security risks, and safely scale enterprise AI.
Built natively into the unified New Relic platform, AI Evaluation goes beyond isolated prompt checks to analyze AI performance across the entire application lifecycle—down to the underlying transaction. By uniting full context rather than focusing on a single metric, it embeds quality and security guardrails directly into existing application performance monitoring workflows.
Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime
Also Read: AI Autonomy Race: How Advanced are Top Countries' AI Strategies?
New Relic Chief Product Officer Brian Emerson says, "Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime. Unlike point solutions that evaluate isolated LLM calls, we look holistically across the entire application transaction to show exactly what happens whenever AI is involved, giving teams full visibility into technical health, business impact, and performance—all within the platform tools that SREs, platform engineers, and developers already use.”



