New Relic launches AI Evaluation for transaction-level AI observability
New Relic introduced AI Evaluation, a framework within its AI Observability platform that scores AI quality and behavior at the transaction level rather than isolated LLM calls. It uses an asynchronous LLM-as-a-judge service, attaches quality scores to distributed traces, and runs guardrail checks for prompt injections, PII leaks and hallucinations.
- Evaluation works at transaction level, not isolated LLM calls
- Asynchronous LLM-as-a-judge service scans live telemetry
- Quality scores are attached to distributed traces
- Guardrails flag prompt injections, jailbreaks and PII leaks
Read next
Software