Anthropic flags gaps in AI guardrails as models grow more capable: Details4Photo© oneindia.com

Anthropic flags gaps in AI guardrails as models grow more capable: Details

, 5 news, 0 views

A claim like "94%" on a tender document looks objective and safe to sign.

Behind a growing share of those numbers sits a practise most buyers have never heard of. When a vendor's system answers ten thousand test questions, somebody has to decide which answers were good. Increasingly, that somebody is not a person but another AI: a large language model, prompted to score the outputs of the system under test.

The industry calls this "Large Language Model (LLM)-as-a-judge." It is fast, cheap, and scales to volumes no human review team can touch. The appeal is real: human evaluation at a national scale is genuinely slow and expensive, which is why LLM-as-a-judge has become standard in corporate evaluation pipelines and on the public leaderboards that rank the world's leading models.

Model judges fail in ways the number does not reveal. The Berkeley-led team that formalised the method in 2023 documented its flaws at the same time: judges favour longer answers, and favour whichever response happens to appear first.