Weighing mainstream and alternative accounts…
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Deeper threads worth pulling on next.
Investigated
Image: frontiersin.org
Yes. Agents can expose sensitive information through final outputs, internal reasoning, shared memory, tool arguments, or outbound tool calls—even during benign tasks or when manipulated by untrusted content. The central disagreement is less whether leakage is possible than how often it occurs in real deployments and whether model-level safeguards are enough without restricting access and data flows.
Two lenses on the same evidence. Source weight and the primary source ratio show what each rests on.
Lens adapted to this topic: Structural and operational critiques
A more skeptical security reading argues that leakage is not merely an occasional model mistake. Agent architectures combine private-data access, exposure to untrusted content, and outbound communication, so probabilistic instructions cannot guarantee prevention. This perspective highlights real-world operational pathways—such as fetched content, auto-loaded resources, exposed keys, and tool wiring—and argues that capability and data-flow restrictions are more dependable than stronger prompts alone.
Deeper threads worth pulling on next.