AI can cite the right document and still give the wrong answer: How businesses are facing a new AI risk

As generative AI evolves, businesses are grappling with a significant risk: the misinterpretation of evidence despite accurate document retrieval. These retrieval-augmented systems can efficiently locate relevant information yet might still draw i...

For years, one of the biggest concerns around generative AI was the possibility of a chatbot simply making things up. Businesses responded by looking for ways to give AI access to reliable information, including company databases, internal documents and other trusted sources. That led to the rapid adoption of retrieval-augmented generation, or RAG.

The idea sounds straightforward. Instead of relying only on what a large language model learned during training, the system first searches a company's information and then uses the retrieved material to formulate an answer. The approach can make AI more useful for everything from document searches and research to legal work and financial analysis.

But there is a problem that businesses are now having to confront: an AI system can find the right document, cite a legitimate source and still reach a conclusion that the source does not actually support. That creates a different kind of AI risk. The problem is no longer only whether the machine can find information. It is whether the answer it produces can be defended against the evidence it cites.


Finding the right document is only the first step

RAG has become an important part of enterprise AI because it allows companies to connect language models with their own information. An employee could, for example, ask an AI system to examine hundreds of internal reports and explain what changed in a particular quarter. A legal team could ask questions about a large collection of case documents. A research team could use AI to search through technical papers.

The technology can dramatically reduce the amount of time spent looking for information. But retrieval and verification are two different things. An AI model may retrieve a document containing relevant information and then misunderstand the context, combine separate facts or make a broader claim than the source permits. Because the final answer is written in fluent language, the mistake may not be obvious to the person reading it.
ADVERTISEMENT

That makes the presence of a citation potentially misleading. A citation tells the user where the AI looked. It does not necessarily prove that the cited material supports every part of the answer.

"Retrieving the right document and generating a convincing answer are not the same as proving that the answer is supported by the source. That distinction becomes critical when AI is being used with research papers, legal documents, financial information or other professional material," said Mayank Ravishankara, a software engineer at Everlaw and an AI researcher.

He said the next step for RAG systems is stronger verification, in which individual claims are checked against the evidence before they reach the user. "We need to move from systems that simply provide an answer with a citation to systems that can establish whether the citation actually supports what is being claimed," Ravishankara said.

Why a citation does not automatically mean an answer is correct
ADVERTISEMENT

Consider a simple example. A company asks its internal AI assistant whether a particular contract contains a limit on damages. The system finds the relevant contract and produces an answer saying that a cap exists. The citation is real. The document is real. The question is whether the document actually says what the AI claims it says.

The model could have confused two different clauses, missed a qualification elsewhere in the agreement or interpreted a general provision as applying to a specific situation.For an employee asking about a routine internal matter, the mistake may be caught quickly. In a legal, financial or compliance workflow, however, the consequences can be much larger.
ADVERTISEMENT

This is why the distinction between grounding and verification is becoming increasingly important. Grounding attempts to keep the AI's answer connected to external evidence. Verification goes one step further by asking whether the evidence actually supports the specific claim.

The problem gets bigger as companies move AI into high-stakes work

The stakes change when AI moves beyond drafting emails or generating marketing ideas. Businesses are increasingly experimenting with AI for contracts, financial research, internal knowledge management, compliance, customer records and other professional tasks. In these situations, an answer that sounds plausible is not enough.

A financial analyst needs to know whether a number came from the correct filing and whether the surrounding context changes its meaning. A lawyer needs to know whether a cited passage actually establishes the proposition being made. A compliance team needs to know whether the underlying rule applies to the particular situation.

In each case, the question is not simply: "Where did the AI get this information?"

It is: "Does this information actually support what the AI is saying?"

That is a much harder problem.

The new enterprise AI metric could be evidence

The first phase of enterprise AI adoption was largely about capability. Could an AI system search a large database? Could it summarise a document? Could it answer questions about a company's internal information? The next phase could be about reliability.

Can the system identify what it does not know? Can it distinguish between information directly supported by a document and an inference made from that information? Can it identify conflicting evidence? And can a user quickly trace an important statement back to the exact material supporting it? That last point is important.

The goal of enterprise AI may not be to eliminate human review altogether. It may instead be to make human review faster and more targeted.

Businesses could face a new productivity paradox

Generative AI was supposed to reduce the amount of time employees spend researching information. But if workers have to manually check every statement produced by an AI assistant against the original documents, some of those productivity gains can disappear. This creates a new challenge for companies deploying AI.

A system that takes 30 seconds to produce an answer but requires 20 minutes of manual verification may not deliver the productivity improvement promised by the technology. The answer, therefore, may lie in building verification into the AI system itself. Instead of giving users a long response followed by a list of sources, future systems could identify individual claims, connect them to supporting passages and flag statements for which sufficient evidence cannot be found.

That would make the AI's reasoning process more transparent without necessarily requiring users to repeat the entire research exercise themselves.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › Industry › Services › Education › AI can cite the right document and still give the wrong answer: How businesses are facing a new AI risk
Text Size:AAA
Success
This article has been saved

*

+