SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
3.40T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2607.25255.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummarySafeFlow is a defense framework for multi-agent LLM systems that treats malicious cross-agent propagation as a semantic information-flow problem. It attaches structured semantic taints to root requests, propagates them through a collaboration graph, and validates workflow-level risk before irreversible actions. Evaluated on four benchmarks covering prompt injection, jailbreak-based tool use, risky code execution, and harmful web-agent behavior, it reduces attack success rates while preserving benign task completion.
Why it mattersFrames multi-agent safety as a propagation problem rather than single-turn classification, with a concrete framework and benchmark results. Useful for anyone designing or auditing delegated agent workflows.
Cited by
No citations on record.
